提交 · 4db40afcfb4dc0a2dd26812a6bfbfd0dc2415187 · PaddlePaddle / Paddle-Lite

18 11月, 2019 5 次提交

add var_conv_2d cuda kernel and unit test test=develop (#2441) · 4db40afc

由 Wilber 提交于 11月 18, 2019

- add var_conv_2d cuda kernel

- add var_conv_2d cuda kernel unit test

- temporarily set to two input mode, remove input(ROW) and input(COLUMN)

4db40afc

[LITE][OPENCL] Enable full and light api for OpenCL (#2331) · cfa086e9

由 Yuan Shuai 提交于 11月 18, 2019

* Fix bug target for kHost and kARM not equal. test=develop

* Fix license. test=develop

* add debug -g option. test=develop

* enable opencl demo. test=develop

* Fix model_optimize_tool found no opencl kernel. test=develop

* add more vlog. test=develop

* remove macro LITE_WITH_OPENCL, LITE_WITH_FPGA in passes. test=develop

* Fix valid_places in mobilenetv1_test. test=develop

* Fix bug of find no real output of fetch, after tool OPs of optimzer passes. test=develop

* Fix vlog as log message in model_optimize_tool. test=develop

* fix miscs. test=develop

* fix comment. test=develop

* Fix misspell of opencl, fpga kernels name in lite/api/CMakeLists.txt. test=develop

* add opencl macro in full_api of demo. test=develop

cfa086e9

P
add sequence_pool cuda kernel, test=develop (#2430) · 3c6b6d2d
由 Pei Yang 提交于 11月 18, 2019
```
add sequence_pool cuda kernel
```
3c6b6d2d
P
add search_group_padding op and x86 kernel, test=develop (#2440) · 33d0cbcd
由 Pei Yang 提交于 11月 18, 2019
```
add search_group_padding op and x86 kernel
```
33d0cbcd

[X86][CUDA] add sequence_arithmetic op , x86 kernel, cuda kernel and unit test (#2436) · 73296cb6

由 zhupengyang 提交于 11月 18, 2019

* [X86][CUDA] add sequence_arithmetic op , x86 kernel, cuda kernel and unit test

test=develop

* add sequence_arithmetic cuda kernel unit test

test=develop

73296cb6

17 11月, 2019 1 次提交
- J
  Add cuda match_matrix_tensor op and test (#2434) · 30a2d0ec
  由 juncaipeng 提交于 11月 17, 2019
```
* add cuda match_matrix_tensor op and test, test=develop
```
  30a2d0ec
16 11月, 2019 1 次提交
- H
  
  [LITE][X86] Add search_aligned_mat_mul and search_seq_fc op for X86 (#2428) · 2148bf49
  由 hong19860320 提交于 11月 16, 2019
  
  2148bf49
15 11月, 2019 2 次提交
- J
  Add content-dnn ops (#2429) · 7f408ee8
  由 juncaipeng 提交于 11月 15, 2019
```
* add search_seq_depadding x86 and cuda
* add match_matrix_tensor x86
* add search_grnn x86, no test
```
  7f408ee8
- L
  fix the cuda bilinear and nearest precision, test=develop (#2426) · cd5d97e3
  由 liu zhengxi 提交于 11月 15, 2019
```
fix the cuda bilinear and nearest precision caused by data type conversion. 
```
  cd5d97e3
14 11月, 2019 5 次提交
- W
  add var_conv_2d op, x86 kernel and unit test test=develop (#2422) · 74f0c8cc
  由 Wilber 提交于 11月 14, 2019
```
- add var_conv_2d op

- add var_conv_2d x86 kernel

- add var_conv_2d x86 test
```
  74f0c8cc
- L
  Fix the compile error for cuda and fix unit test (#2424) · 4473ef43
  由 liu zhengxi 提交于 11月 14, 2019
```
* fix the compile error for cuda and fix unit test
```
  4473ef43
- L
  fix conv2d kernel bugs that results in precision diff test=develop (#2420) · 8b36f2aa
  由 lijianshe02 提交于 11月 14, 2019
```
* fix conv kernel bugs and open mobilenet ci test=develop
```
  8b36f2aa
- H
  
  [LITE][NPU] Upgrade HiAI DDK from 300 to 310 (#2423) · 553f314d
  由 hong19860320 提交于 11月 14, 2019
  
  553f314d
- P
  Update lookup_table op on arm x86，add lookup_table_v2_op (#2405) · 24029728
  由 Pei Yang 提交于 11月 14, 2019
```
* update lookup_table arm x86, test=develop

* add lookup_table_v2_op for compatibility, test=develop
```
  24029728
13 11月, 2019 4 次提交
- W
  add sequence_reverse op and kerenl for arm and cuda test=develop (#2397) · 3132ad03
  由 Wilber 提交于 11月 13, 2019
```
- add sequence_reverse op

- add sequence_reverse kernel for x86 and cuda

- add sequence_reverse_test for x86 and cuda
```
  3132ad03
- L
  Add cast op for x86 platform on Paddle-Lite (#2413) · 7a32c431
  由 liu zhengxi 提交于 11月 13, 2019
```
* add cast op for x86 platform

* alter the struct to class to hide the data and alter the pointer
```
  7a32c431
- W
  add sequence_concat op kernel and test test=develop (#2414) · 46bb9703
  由 Wilber 提交于 11月 13, 2019
```
- add sequence_concat op

- add sequence_concat kernel for x86 and cuda

- add sequence_concat_test for x86 and cuda
```
  46bb9703
- L
  Update the ops to fluid (#2406) · c77e4074
  由 liu zhengxi 提交于 11月 13, 2019
```
align the lite nearest， bilinear op to fluid on arm and cuda
```
  c77e4074
12 11月, 2019 1 次提交
- J
  Upgrade concat and unsqueeze, test=develop (#2378) · 284e8166
  由 juncaipeng 提交于 11月 12, 2019
```
* update concat and unsqueeze, test=develop
```
  284e8166
11 11月, 2019 2 次提交

P
add cuda kernel:lookup table, test=develop (#2403) · 0724abba
由 Pei Yang 提交于 11月 11, 2019
```
add cuda kernel:lookup table
```
0724abba

fix pool bug and speed, test=develop (#2385) · f85c3689

由 HappyAngel 提交于 11月 11, 2019

* fix pooling bug and speed

* fix build error

* delete VLOG in pool, test=develop

* add openmp, test=develop

* fix lite/kernels/arm/pool_compute_test basic_pooling compute error bug, test=develop

f85c3689

08 11月, 2019 2 次提交

Move muliti class kernel back to basic (#2396) · 7ea34b1b

由 huzhiqiang 提交于 11月 07, 2019

* move multiclass_nms kernel back to host test=develop

* move layer_norm OP and arm_kernel into extra type since it's added after release/v2.0-beta1 and not related with CV test=develop

* fix code_style test=develop

7ea34b1b

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the... · e7610023

由 hong19860320 提交于 11月 08, 2019

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the limitation of the dimensions of input tensors in the graph compute kernel, and refine log message (#2395)

test=develop

e7610023

07 11月, 2019 2 次提交

T
fix:fix deviceinfo worksapce to tls · 30300c98
由 TianXiaogang 提交于 11月 07, 2019
```
mv deviceinfo.workspace and other relative member to thread_local_storage
```
30300c98

check arm kernels type to make sure all_library_links work normally (#2386) · 6b38eab8

由 huzhiqiang 提交于 11月 06, 2019

We have changed 11 arm_kernels into extra type in #2347 , which has caused test_compiling failure. In this PR , we move their 11 related arm_kernel_test into build_extra=ON

6b38eab8

06 11月, 2019 3 次提交

update slice and reshape op and test on one op fake model test=develop (#2377) · 126964a1

由 Wilber 提交于 11月 06, 2019

update reshape op to support multiple input types of shape.
priority: input(ShapeTensor) > input(Shape) > attr(shape)

update slice op to support multiple iput types of starts and ends.
priority: input(StartsTensor) > input(StartsTensorList) > attr(starts)

126964a1

fix fill_constant kernel bug test=develop (#2376) · f607870c

由 Wilber 提交于 11月 06, 2019

fill_constant kernel only registered float type, only the float data type is produced, which is obviously a bug.

Now, produce data based on the data type attr.

By the way, fix the cast kernel bug.

f607870c

H

change arm conv_kernels into basic to support mobilenetv1 test=develop (#2375) · 62bbf501
由 huzhiqiang 提交于 11月 06, 2019

62bbf501

05 11月, 2019 1 次提交
- L
  fix StepRNN model run related bugs (#2300) · 99d4f70e
  由 lijianshe02 提交于 11月 05, 2019
```
* fix step rnn model run bugs test=develop
```
  99d4f70e
04 11月, 2019 1 次提交
- H
  Move new op kernel into extra (#2348) · a438d0dc
  由 huzhiqiang 提交于 11月 04, 2019
```
* move some basic ops into extra type to reduce library size test=develop (#2347)
```
  a438d0dc
01 11月, 2019 2 次提交
- Z
  [XPU] add batchnorm op bridge and unit test (#2323) · b65861ee
  由 zhupengyang 提交于 11月 01, 2019
```
* [XPU] add batchnorm op bridge and unit test

test=develop

* fix DMLC_USE_GLOG

test=develop
```
  b65861ee
- T
  
  fix: fix conv_direct && test=develop (#2314) · ecca7325
  由 TianXiaogang 提交于 11月 01, 2019
  
  ecca7325
30 10月, 2019 2 次提交
- Z
  [XPU] update elemetwise_add, conv and mul ops (#2293) · 2a1856a9
  由 zhupengyang 提交于 10月 30, 2019
```
test=develop
```
  2a1856a9
- H
  
  [LITE][NPU] Use FullConnection op to solve the compatibility between Kirin 810 and 990 (#2283) · a5fe6f8e
  由 hong19860320 提交于 10月 30, 2019
  
  a5fe6f8e
29 10月, 2019 5 次提交
- L
  Add tanh op and gelu op for x86 platform (#2265) · e3368aa4
  由 liu zhengxi 提交于 10月 29, 2019
```
* add tanh op in x86 platform and its unittest, test=develop

* add gelu op on x86 platform and add its unittests, test=develop

* update depends for math_function for activation for gelu, test=develop
```
  e3368aa4
- Z
  [XPU] fix elementwise op bridge when x or y are from weight (#2272) · 40baeeae
  由 zhupengyang 提交于 10月 29, 2019
```
test=develop
```
  40baeeae
- H
  [LITE][NPU] Add supporting for Huawei offical DDK (#2262) · 92f0b3c5
  由 hong19860320 提交于 10月 29, 2019
```
* Add supporting for Huawei offical DDK
* Fix the param of graph op in NPU graph computing kernel
```
  92f0b3c5
- G
  modify layer_norm_compute_test.cc (#2261) · f2b9ca34
  由 guofei 提交于 10月 29, 2019
```
test=develop
```
  f2b9ca34
- Z
  [XPU] add mul op bridge (#2267) · 941c487d
  由 zhupengyang 提交于 10月 29, 2019
```
* [XPU] add mul op bridge and unit test

test=develop

* use tmp tensor for transposed y

test=develop
```
  941c487d
28 10月, 2019 1 次提交
- Z
  [XPU] add elementwise, pool, softmax op bridges and unit tests (#2264) · 231a658b
  由 zhupengyang 提交于 10月 28, 2019
```
test=develop
```
  231a658b