提交 · a7f7d49bddfd1fa805d6d00ebe4c515f99141f58 · PaddlePaddle / Paddle-Lite

22 11月, 2019 2 次提交

update pooling 2-padding to 4-padding (#2410) · a7f7d49b

由 HappyAngel 提交于 11月 22, 2019

* fix pooling bug and speed

* fix build error

* delete VLOGin pool, test=develop

* add openmp, test=develop

* fix lite/kernels/arm/pool_compute_test basic_pooling compute error bug, test=develop

* update pooling 2-pad to 4-pad, test=develop

* fix 2-pad to 4-pad in operators/pool_op.h, AttachKernel will set param, so 2-pad to 4-pad funcs should put in AttachKernel. test=ddevellop

* put 2-pad to 4-pad in AttachImpl, test=develop

* according to reviews, fix some format error. test=develop

* fix format errorr, add (). test=develop

* change paddings type to support dynamically modify, test=develop

* update padding type int other devices, test=develop

* fix x8d build error on shared_ptr, test=ddevelop

* fix formmat in operators pool_op.cc, test=develop

a7f7d49b

P

add search_group_padding cuda kernel, test=develop (#2472) · 36c0068e
由 Pei Yang 提交于 11月 22, 2019

36c0068e

21 11月, 2019 3 次提交
- H
  add cuda kernel for sequence_topk_avg_pooling and search_fc (#2451) · 3a881861
  由 huzhiqiang 提交于 11月 21, 2019
```
* cuda kernel for sequence_topk_avg_pooling and search_fc test=develop
```
  3a881861
- P
  
  remove duplicate cmake targets of sequence-pool (#2467) · bf2c6fca
  由 Pei Yang 提交于 11月 21, 2019
  
  bf2c6fca
- 石
  fix cuda build error, test=develop (#2464) · d8ddbcc6
  由石晓伟提交于 11月 21, 2019
```
* fix cuda building, test=develop

* remove sequence_pool from cmake because build error, test=develop
```
  d8ddbcc6
20 11月, 2019 7 次提交
- P
  fix sequence pool cuda (#2466) · 43f1358f
  由 Pei Yang 提交于 11月 20, 2019
```
* add sequence_pool cuda kernel, test=develop

* fix sequence_pool cuda,test=develop

* fix and complete unittest, test=develop

* fix macro of sequence_pool cuda, test=develop
```
  43f1358f
- J
  fix x86 search_grnn, add cuda search_grnn and unit test (#2448) · e1b67433
  由 juncaipeng 提交于 11月 20, 2019
```
* fix x86 search_grnn and add unit test
* add cuda search_grnn and unit test
```
  e1b67433
- L
  
  Add layer_norm op on Lite x86 platform (#2463) · fa8c8971
  由 liu zhengxi 提交于 11月 20, 2019
  
  fa8c8971
- L
  Add stack op on Lite x86 platform and fix extra cmake error (#2458) · 79f8f42d
  由 liu zhengxi 提交于 11月 20, 2019
```
* add stack op and its unit tests, test=develop
```
  79f8f42d
- Y
  [ARM] sgemv support transA, test=develop (#2453) · dde12f0d
  由 yiicy 提交于 11月 20, 2019
```
* [ARM] sgemv support transA, test=develop

* add sgemv ut, test=develop
```
  dde12f0d
- P
  fix sequence pool cuda (#2457) · b094b2b6
  由 Pei Yang 提交于 11月 20, 2019
```
* add sequence_pool cuda kernel, test=develop

* fix sequence_pool cuda,test=develop

* fix and complete unittest, test=develop
```
  b094b2b6
- 石
  add cuda ci building, test=develop (#2460) · 2621af0e
  由石晓伟提交于 11月 20, 2019
```
* add cuda ci building, test=develop

* update comments, test=develop
```
  2621af0e
19 11月, 2019 5 次提交
- Y
  
  fix lrn param, align to fluid, test=develop (#2452) · 94255f6c
  由 yiicy 提交于 11月 19, 2019
  
  94255f6c
- H
  
  [LITE][CUDA] Add CUDA kernel for search_aligned_mat_mul and search_seq_fc Op (#2449) · 8373aec5
  由 hong19860320 提交于 11月 19, 2019
  
  8373aec5
- Z
  add search_seq_softmax op; regist search_seq_softmax x86 kernel and cuda kernel (#2445) · f9930fc1
  由 zhupengyang 提交于 11月 19, 2019
```
test=develop
```
  f9930fc1
- H
  add x86 kernels: search_fc and sequence_topk_ave_pooling (#2443) · 68fe5b5c
  由 huzhiqiang 提交于 11月 18, 2019
```
* add x86 op and kernel : search_fc and sequence_topk_avg_pooling   for content-dnn model test=develop
```
  68fe5b5c
- Z
  [X86][CUDA] add attention_padding_mask op, x86 kernel, cuda kernel and unit tests (#2437) · ef6f7b84
  由 zhupengyang 提交于 11月 19, 2019
```
* [X86] add attention_padding_mask op, x86 kernel and unit test

test=develop

* [CUDA] add attention_padding_mask cuda kernel and unit test

test=develop
```
  ef6f7b84
18 11月, 2019 5 次提交

add var_conv_2d cuda kernel and unit test test=develop (#2441) · 884c840d

由 Wilber 提交于 11月 18, 2019

- add var_conv_2d cuda kernel

- add var_conv_2d cuda kernel unit test

- temporarily set to two input mode, remove input(ROW) and input(COLUMN)

884c840d

[LITE][OPENCL] Enable full and light api for OpenCL (#2331) · d242bdfb

由 Yuan Shuai 提交于 11月 18, 2019

* Fix bug target for kHost and kARM not equal. test=develop

* Fix license. test=develop

* add debug -g option. test=develop

* enable opencl demo. test=develop

* Fix model_optimize_tool found no opencl kernel. test=develop

* add more vlog. test=develop

* remove macro LITE_WITH_OPENCL, LITE_WITH_FPGA in passes. test=develop

* Fix valid_places in mobilenetv1_test. test=develop

* Fix bug of find no real output of fetch, after tool OPs of optimzer passes. test=develop

* Fix vlog as log message in model_optimize_tool. test=develop

* fix miscs. test=develop

* fix comment. test=develop

* Fix misspell of opencl, fpga kernels name in lite/api/CMakeLists.txt. test=develop

* add opencl macro in full_api of demo. test=develop

d242bdfb

P
add sequence_pool cuda kernel, test=develop (#2430) · 3d73dea9
由 Pei Yang 提交于 11月 18, 2019
```
add sequence_pool cuda kernel
```
3d73dea9
P
add search_group_padding op and x86 kernel, test=develop (#2440) · 1e88d1e8
由 Pei Yang 提交于 11月 18, 2019
```
add search_group_padding op and x86 kernel
```
1e88d1e8

[X86][CUDA] add sequence_arithmetic op , x86 kernel, cuda kernel and unit test (#2436) · 8599c042

由 zhupengyang 提交于 11月 18, 2019

* [X86][CUDA] add sequence_arithmetic op , x86 kernel, cuda kernel and unit test

test=develop

* add sequence_arithmetic cuda kernel unit test

test=develop

8599c042

17 11月, 2019 1 次提交
- J
  Add cuda match_matrix_tensor op and test (#2434) · ce21ff5d
  由 juncaipeng 提交于 11月 17, 2019
```
* add cuda match_matrix_tensor op and test, test=develop
```
  ce21ff5d
16 11月, 2019 1 次提交
- H
  
  [LITE][X86] Add search_aligned_mat_mul and search_seq_fc op for X86 (#2428) · 78f76834
  由 hong19860320 提交于 11月 16, 2019
  
  78f76834
15 11月, 2019 2 次提交
- J
  Add content-dnn ops (#2429) · 603b810f
  由 juncaipeng 提交于 11月 15, 2019
```
* add search_seq_depadding x86 and cuda
* add match_matrix_tensor x86
* add search_grnn x86, no test
```
  603b810f
- L
  fix the cuda bilinear and nearest precision, test=develop (#2426) · 76766636
  由 liu zhengxi 提交于 11月 15, 2019
```
fix the cuda bilinear and nearest precision caused by data type conversion. 
```
  76766636
14 11月, 2019 5 次提交
- W
  add var_conv_2d op, x86 kernel and unit test test=develop (#2422) · 9dcd9914
  由 Wilber 提交于 11月 14, 2019
```
- add var_conv_2d op

- add var_conv_2d x86 kernel

- add var_conv_2d x86 test
```
  9dcd9914
- L
  Fix the compile error for cuda and fix unit test (#2424) · 1f075a8b
  由 liu zhengxi 提交于 11月 14, 2019
```
* fix the compile error for cuda and fix unit test
```
  1f075a8b
- L
  fix conv2d kernel bugs that results in precision diff test=develop (#2420) · bfd2a950
  由 lijianshe02 提交于 11月 14, 2019
```
* fix conv kernel bugs and open mobilenet ci test=develop
```
  bfd2a950
- H
  
  [LITE][NPU] Upgrade HiAI DDK from 300 to 310 (#2423) · c1837d76
  由 hong19860320 提交于 11月 14, 2019
  
  c1837d76
- P
  Update lookup_table op on arm x86，add lookup_table_v2_op (#2405) · 94731268
  由 Pei Yang 提交于 11月 14, 2019
```
* update lookup_table arm x86, test=develop

* add lookup_table_v2_op for compatibility, test=develop
```
  94731268
13 11月, 2019 4 次提交
- W
  add sequence_reverse op and kerenl for arm and cuda test=develop (#2397) · acf09294
  由 Wilber 提交于 11月 13, 2019
```
- add sequence_reverse op

- add sequence_reverse kernel for x86 and cuda

- add sequence_reverse_test for x86 and cuda
```
  acf09294
- L
  Add cast op for x86 platform on Paddle-Lite (#2413) · 5bcbacb7
  由 liu zhengxi 提交于 11月 13, 2019
```
* add cast op for x86 platform

* alter the struct to class to hide the data and alter the pointer
```
  5bcbacb7
- W
  add sequence_concat op kernel and test test=develop (#2414) · 8a1d942a
  由 Wilber 提交于 11月 13, 2019
```
- add sequence_concat op

- add sequence_concat kernel for x86 and cuda

- add sequence_concat_test for x86 and cuda
```
  8a1d942a
- L
  Update the ops to fluid (#2406) · 518a87ef
  由 liu zhengxi 提交于 11月 13, 2019
```
align the lite nearest， bilinear op to fluid on arm and cuda
```
  518a87ef
12 11月, 2019 1 次提交
- J
  Upgrade concat and unsqueeze, test=develop (#2378) · 26470600
  由 juncaipeng 提交于 11月 12, 2019
```
* update concat and unsqueeze, test=develop
```
  26470600
11 11月, 2019 2 次提交

P
add cuda kernel:lookup table, test=develop (#2403) · 15eccb9e
由 Pei Yang 提交于 11月 11, 2019
```
add cuda kernel:lookup table
```
15eccb9e

fix pool bug and speed, test=develop (#2385) · d197de00

由 HappyAngel 提交于 11月 11, 2019

* fix pooling bug and speed

* fix build error

* delete VLOG in pool, test=develop

* add openmp, test=develop

* fix lite/kernels/arm/pool_compute_test basic_pooling compute error bug, test=develop

d197de00

08 11月, 2019 2 次提交

Move muliti class kernel back to basic (#2396) · 52e0db46

由 huzhiqiang 提交于 11月 07, 2019

* move multiclass_nms kernel back to host test=develop

* move layer_norm OP and arm_kernel into extra type since it's added after release/v2.0-beta1 and not related with CV test=develop

* fix code_style test=develop

52e0db46

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the... · aacba6f5

由 hong19860320 提交于 11月 08, 2019

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the limitation of the dimensions of input tensors in the graph compute kernel, and refine log message (#2395)

test=develop

aacba6f5