提交 · 90e5895aa49347c17c33d67c440293db214f68e1 · PaddlePaddle / Paddle-Lite

16 12月, 2019 3 次提交

T
update fpga KD patch (#2609) · 90e5895a
由 TianXiaogang 提交于 12月 16, 2019
```
* fix: update backend fpga patch
```
90e5895a

[LITE][OPENCL] Add relu image2d kernel unit test, Fix conv2d_1x1, relu, layout... · ee177c6b

由 Yuan Shuai 提交于 12月 16, 2019

[LITE][OPENCL] Add relu image2d kernel unit test, Fix conv2d_1x1, relu, layout using new Image2D Layout (#2564)

* add 3 layout for opencl image. test=develop

* add relu image2d test. test=develop

ee177c6b

[LITE][OPENCL] Add depthwise_conv_3x3 opencl kernel (#2601) · 600c8c20

由 Jiaying Zhao 提交于 12月 16, 2019

* [LITE][OPENCL] Add depthwise_conv_3x3 opencl kernel

* [LITE][OPENCL] Add depthwise_conv_3x3 opencl kernel. test=develop

* [LITE][OPENCL] Add Pool opencl kernel. test=develop

600c8c20

15 12月, 2019 1 次提交
- W
  optimize search_grnn test=develop (#2608) · dad43f81
  由 Wilber 提交于 12月 15, 2019
```
optimize search_grnn
```
  dad43f81
13 12月, 2019 1 次提交
- H
  [LITE][NPU][XPU] Refine subgraph pass, and support NPU/XPU model generation at... · d5434aa2
  由 hong19860320 提交于 12月 13, 2019
```
[LITE][NPU][XPU] Refine subgraph pass, and support NPU/XPU model generation at execution time (#2576)
```
  d5434aa2
12 12月, 2019 1 次提交

[LITE][OPENCL] Add conv2d_1x1 opencl kernel (#2591) · e583a55d

由 xiebaiyuan 提交于 12月 12, 2019

* add opencl conv1x1 image impl and unit test pass with relu & bias,
add layout_compute --> buffer2image float32 --> with unit test pass
suite checked test for more situation , test=develop

* add opencl conv1x1 image impl and unit test pass with relu & bias,
add layout_compute --> buffer2image float32 --> with unit test pass
suite checked test for more situation , test=develop

* fix white space cpp lint , test=develop

e583a55d

11 12月, 2019 1 次提交
- T
  
  add winograd f23 implement (#2584) · f99c34c8
  由 TianXiaogang 提交于 12月 11, 2019
  
  f99c34c8
10 12月, 2019 1 次提交
- Y
  
  [ARM] add instance norm op and ut, test=develop (#2578) · 9a3552db
  由 yiicy 提交于 12月 10, 2019
  
  9a3552db
09 12月, 2019 2 次提交
- Y
  
  fix ios demo build error, test=develop (#2579) · 0c44ac9c
  由 yiicy 提交于 12月 09, 2019
  
  0c44ac9c
- Z
  [NPU] support relu6 (#2582) · a2f981a4
  由 zhupengyang 提交于 12月 09, 2019
```
test=develop
```
  a2f981a4
07 12月, 2019 1 次提交

Support mask_rcnn (#2484) · c2f72cb3

由 juncaipeng 提交于 12月 07, 2019

* add arm split lod tensor, test=develop

* add arm merge lod tensor, test=develop

* update split merge lod tensor, test=develop

* add reduce_prob op, test=develop

* support mask_rcnn succeed, test=develop

c2f72cb3

04 12月, 2019 1 次提交

[cuda] [int8] resnet50 cuda int8 support (#2417) · f7574646

由 Zhaolong Xing 提交于 12月 04, 2019

* init resnet cuda int8 support
test=develop

* refine cuda unit test
test=develop

* add the forgeted file.
test=develop

f7574646

03 12月, 2019 1 次提交
- T
  Armv8 4x4 gemm (#2528) · 1ebac1c0
  由 TianXiaogang 提交于 12月 03, 2019
```
* feat: add sgemm4x4 for armv8

* fix: fix armv7 gemm choose condition
```
  1ebac1c0
30 11月, 2019 1 次提交
- Z
  [NPU] fix act; refine act unit tests; fix batch_norm (#2533) · 0349bfd0
  由 zhupengyang 提交于 11月 30, 2019
```
test=develop
```
  0349bfd0
28 11月, 2019 1 次提交
- Y
  
  [cherry-pick][ARM] conv_transpose operator support padding_algorithm, test=develop (#2500) · 5fac0949
  由 yiicy 提交于 11月 28, 2019
  
  5fac0949
27 11月, 2019 1 次提交
- H
  
  add into beam_search.cc test=develop (#2506) · e8ea4a56
  由 huzhiqiang 提交于 11月 27, 2019
  
  e8ea4a56
26 11月, 2019 1 次提交

add winograd c4 implement (#2494) · e0eee83c

由 TianXiaogang 提交于 11月 26, 2019

fix: fix conv_block prepack_input_nxwc4 bug
* fix: optimize sgemm_c4 in armv7
     change condition of choose winograd kernel
* fix: change conv choose kernel condition

e0eee83c

25 11月, 2019 1 次提交

[cherry-pick][ARM] armv7 improve sgemmc4 small kernel speed by add 4x8 block, test=develop (#2486) · e909ffa4

由 yiicy 提交于 11月 25, 2019

* unfinish sgemmc4

* finish armv8 sgemmc4

* arm add sgemmc4 with deal with remain

* [ARM] add sgemmc4 small kernel, test=develop

* [ARM] sgemmc4 small improve armv7 speed by add 4x8 block, test=develop

e909ffa4

22 11月, 2019 4 次提交

update conv 2-pad to 4-pad (#2404) · 820eb6d4

由 HappyAngel 提交于 11月 22, 2019

* fix conv 2-pad to 4-pad

* fix compute conv shape

* fix pad, test=develop

* change conv_depthwise_3x3s1_fp.cc name to conv3x3s1p01_depthwise_fp32.cc to distinguish between conv3x3s1_depthwise_fp32.cc

* delete printf note in conv3x3s1, test=develop

* delete printf note, test=develop

* delete gem_sdot.h, test=develop

it is coped from __gemm_sdot_meta_.h

* update compute padding, test=develop

* fix padding size, must be 2 or 4. test=develop

* fix format in operators/conv_op.cc, test=develop

* change #if 0 to #if 1, test=develop

* put 2-pad to 4-pad in AttachImpl, test=develop

* fix clang-format error inn tests/math/connv_compute_test, test=develop

* fix x86 test result error, test=develop

* add asymmetric padding test case in liite/tests/math/conv_compute.cc, test=develop

* change paddings type to support dynamically modify, test=develop

* fix x86 build error in connv_compute_test, test=develop

* fix opencl build error, test=develop

* fix oopencl build error, test=develop

* fix  opencl/conv_compute build error, test=develop

* fix  opencl/conv_compute build error, test=develop

* fix format in kernels/opencl/conv_computte_ttest,test=develop

* fix build error, test=develop

fix build error in kernels/x86/conv_compute.h

820eb6d4

add NHWC NCHW transform, test=develop (#2381) · 6b3c341f

由 HappyAngel 提交于 11月 22, 2019

* add nhwc to nchw

* add layout in funcs

* change layout as extra, test=develop

* change make, test=develop

* use template class method to update layout NNCHHW and NHWC transform, test=develop

* fix cmake error, set layout to extra, test=develop

* fix test_layout_compute_arm test, its extra

* layout is extra, test=develop

* fix error in kernels/arm/layout_comput.cc when register kernel, DataLayout must be NCHW, test=develop

* delete extra note, test=develop

* delete extra test

* delete layout_test, test=develop

, its in tests/math/layout_comput_test

* delete extrat test, test=develop

6b3c341f

Y
[ARM] add sgemmc4 common and small kernel, support for winograd, test=develop (#2471) · 66d2ae25
由 yiicy 提交于 11月 22, 2019
```
* unfinish sgemmc4

* finish armv8 sgemmc4

* arm add sgemmc4 with deal with remain

* [ARM] add sgemmc4 small kernel, test=develop
```
66d2ae25

update pooling 2-padding to 4-padding (#2410) · a7f7d49b

由 HappyAngel 提交于 11月 22, 2019

* fix pooling bug and speed

* fix build error

* delete VLOGin pool, test=develop

* add openmp, test=develop

* fix lite/kernels/arm/pool_compute_test basic_pooling compute error bug, test=develop

* update pooling 2-pad to 4-pad, test=develop

* fix 2-pad to 4-pad in operators/pool_op.h, AttachKernel will set param, so 2-pad to 4-pad funcs should put in AttachKernel. test=ddevellop

* put 2-pad to 4-pad in AttachImpl, test=develop

* according to reviews, fix some format error. test=develop

* fix format errorr, add (). test=develop

* change paddings type to support dynamically modify, test=develop

* update padding type int other devices, test=develop

* fix x8d build error on shared_ptr, test=ddevelop

* fix formmat in operators pool_op.cc, test=develop

a7f7d49b

21 11月, 2019 1 次提交

石

fix cuda build error, test=develop (#2464) · d8ddbcc6

由石晓伟提交于 11月 21, 2019

* fix cuda building, test=develop

* remove sequence_pool from cmake because build error, test=develop

d8ddbcc6

20 11月, 2019 1 次提交
- Y
  [ARM] sgemv support transA, test=develop (#2453) · dde12f0d
  由 yiicy 提交于 11月 20, 2019
```
* [ARM] sgemv support transA, test=develop

* add sgemv ut, test=develop
```
  dde12f0d
19 11月, 2019 2 次提交
- H
  
  [LITE][CUDA] Add CUDA kernel for search_aligned_mat_mul and search_seq_fc Op (#2449) · 8373aec5
  由 hong19860320 提交于 11月 19, 2019
  
  8373aec5
- H
  add x86 kernels: search_fc and sequence_topk_ave_pooling (#2443) · 68fe5b5c
  由 huzhiqiang 提交于 11月 18, 2019
```
* add x86 op and kernel : search_fc and sequence_topk_avg_pooling   for content-dnn model test=develop
```
  68fe5b5c
18 11月, 2019 1 次提交

[LITE][OPENCL] Enable full and light api for OpenCL (#2331) · d242bdfb

由 Yuan Shuai 提交于 11月 18, 2019

* Fix bug target for kHost and kARM not equal. test=develop

* Fix license. test=develop

* add debug -g option. test=develop

* enable opencl demo. test=develop

* Fix model_optimize_tool found no opencl kernel. test=develop

* add more vlog. test=develop

* remove macro LITE_WITH_OPENCL, LITE_WITH_FPGA in passes. test=develop

* Fix valid_places in mobilenetv1_test. test=develop

* Fix bug of find no real output of fetch, after tool OPs of optimzer passes. test=develop

* Fix vlog as log message in model_optimize_tool. test=develop

* fix miscs. test=develop

* fix comment. test=develop

* Fix misspell of opencl, fpga kernels name in lite/api/CMakeLists.txt. test=develop

* add opencl macro in full_api of demo. test=develop

d242bdfb

17 11月, 2019 1 次提交
- J
  Add cuda match_matrix_tensor op and test (#2434) · ce21ff5d
  由 juncaipeng 提交于 11月 17, 2019
```
* add cuda match_matrix_tensor op and test, test=develop
```
  ce21ff5d
14 11月, 2019 1 次提交
- H
  
  [LITE][NPU] Upgrade HiAI DDK from 300 to 310 (#2423) · c1837d76
  由 hong19860320 提交于 11月 14, 2019
  
  c1837d76
13 11月, 2019 1 次提交
- L
  Update the ops to fluid (#2406) · 518a87ef
  由 liu zhengxi 提交于 11月 13, 2019
```
align the lite nearest， bilinear op to fluid on arm and cuda
```
  518a87ef
11 11月, 2019 1 次提交

fix pool bug and speed, test=develop (#2385) · d197de00

由 HappyAngel 提交于 11月 11, 2019

* fix pooling bug and speed

* fix build error

* delete VLOG in pool, test=develop

* add openmp, test=develop

* fix lite/kernels/arm/pool_compute_test basic_pooling compute error bug, test=develop

d197de00

08 11月, 2019 1 次提交

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the... · aacba6f5

由 hong19860320 提交于 11月 08, 2019

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the limitation of the dimensions of input tensors in the graph compute kernel, and refine log message (#2395)

test=develop

aacba6f5

07 11月, 2019 1 次提交

fix jit_matmul bug according to paddle pr#20948 test=develop (#2392) · f51c4891

由 Wilber 提交于 11月 07, 2019

fix jit::matmul bug. Input x shape is (m, k), weight shape is (k, n). When k < 512, m==1, and n is a multiple of 16, the weight pointer is not correctly updated in the group calculation in the implementation of jit::matmul, resulting in the result diff

f51c4891

30 10月, 2019 1 次提交
- H
  
  [LITE][NPU] Use FullConnection op to solve the compatibility between Kirin 810 and 990 (#2283) · f3e5d3e5
  由 hong19860320 提交于 10月 30, 2019
  
  f3e5d3e5
29 10月, 2019 2 次提交

Add tanh op and gelu op for x86 platform (#2265) · 605a309c

由 liu zhengxi 提交于 10月 29, 2019

* add tanh op in x86 platform and its unittest, test=develop

* add gelu op on x86 platform and add its unittests, test=develop

* update depends for math_function for activation for gelu, test=develop

605a309c

H
[LITE][NPU] Add supporting for Huawei offical DDK (#2262) · dc2b853e
由 hong19860320 提交于 10月 29, 2019
```
* Add supporting for Huawei offical DDK
* Fix the param of graph op in NPU graph computing kernel
```
dc2b853e

28 10月, 2019 1 次提交

[LITE][XPU] initial support for XPU (#2202) · 06d058fe

由 hong19860320 提交于 10月 28, 2019

* Initial support for XPU
* Fix compiling errors of XPU
* Move XPU op kernel bridges from backends to kernels to fix deps order
* Change the namespace and directory of XPU bridges
* Add XPU SDK
* Fix header files and namespace of XPU SDK
* Add unit tests for relu and conv2d ops
* Restore the modification of paddle_api_test
* Supports simple model which contains only a relu layer
* Add compiling scripts for XPU
* Fix compiling errors of XPU
* Add comments for XPU LoadModel and BuildModel

06d058fe

24 10月, 2019 1 次提交

Fix gemv int8 error (#2249) · 3406a11a

由 Xiaoyang LI 提交于 10月 24, 2019

* remove log in reshape, fix conv error when padding size=4, test=develop

* fix style, test=develop

* remove useless code, test=develop

* remove redundant model test file, test=develop

* change cluster to power_mode, test=develop

* fix build error, test=develop

* change cluster to power_mode, test=develop

* change opt_nb to use_optimize_nb, test=develop

* null, test=develop

* add gemv-int8 test, fix clang build error, test=develop

* fix gemv-int8 error when build with clang, test=develop

3406a11a

23 10月, 2019 2 次提交

W
modify yolobox_cuda to support multiple runs (#2245) · a4a19ba4
由 Wilber 提交于 10月 23, 2019
```
* modify yolobox_cuda to support multiple runs test=develop
```
a4a19ba4

Enable pool2d, dropout, transpose and transpose2 op on x86 (#2226) · aa507f9b

由 liu zhengxi 提交于 10月 23, 2019

* enable pool2d op on x86 and add its unit tests, test=develop

* enable dropout op and add its unit tests, test=develop

* add tranpose, transpose2 op and add their unit tests, test=develop

aa507f9b