提交 · 46bb9703455a60da91da10c1b387cf874c437f3f · PaddlePaddle / Paddle-Lite

13 11月, 2019 2 次提交
- W
  add sequence_concat op kernel and test test=develop (#2414) · 46bb9703
  由 Wilber 提交于 11月 13, 2019
```
- add sequence_concat op

- add sequence_concat kernel for x86 and cuda

- add sequence_concat_test for x86 and cuda
```
  46bb9703
- L
  Update the ops to fluid (#2406) · c77e4074
  由 liu zhengxi 提交于 11月 13, 2019
```
align the lite nearest， bilinear op to fluid on arm and cuda
```
  c77e4074
12 11月, 2019 1 次提交
- J
  Upgrade concat and unsqueeze, test=develop (#2378) · 284e8166
  由 juncaipeng 提交于 11月 12, 2019
```
* update concat and unsqueeze, test=develop
```
  284e8166
11 11月, 2019 2 次提交

P
add cuda kernel:lookup table, test=develop (#2403) · 0724abba
由 Pei Yang 提交于 11月 11, 2019
```
add cuda kernel:lookup table
```
0724abba

fix pool bug and speed, test=develop (#2385) · f85c3689

由 HappyAngel 提交于 11月 11, 2019

* fix pooling bug and speed

* fix build error

* delete VLOG in pool, test=develop

* add openmp, test=develop

* fix lite/kernels/arm/pool_compute_test basic_pooling compute error bug, test=develop

f85c3689

08 11月, 2019 2 次提交

Move muliti class kernel back to basic (#2396) · 7ea34b1b

由 huzhiqiang 提交于 11月 07, 2019

* move multiclass_nms kernel back to host test=develop

* move layer_norm OP and arm_kernel into extra type since it's added after release/v2.0-beta1 and not related with CV test=develop

* fix code_style test=develop

7ea34b1b

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the... · e7610023

由 hong19860320 提交于 11月 08, 2019

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the limitation of the dimensions of input tensors in the graph compute kernel, and refine log message (#2395)

test=develop

e7610023

07 11月, 2019 2 次提交

T
fix:fix deviceinfo worksapce to tls · 30300c98
由 TianXiaogang 提交于 11月 07, 2019
```
mv deviceinfo.workspace and other relative member to thread_local_storage
```
30300c98

check arm kernels type to make sure all_library_links work normally (#2386) · 6b38eab8

由 huzhiqiang 提交于 11月 06, 2019

We have changed 11 arm_kernels into extra type in #2347 , which has caused test_compiling failure. In this PR , we move their 11 related arm_kernel_test into build_extra=ON

6b38eab8

06 11月, 2019 3 次提交

update slice and reshape op and test on one op fake model test=develop (#2377) · 126964a1

由 Wilber 提交于 11月 06, 2019

update reshape op to support multiple input types of shape.
priority: input(ShapeTensor) > input(Shape) > attr(shape)

update slice op to support multiple iput types of starts and ends.
priority: input(StartsTensor) > input(StartsTensorList) > attr(starts)

126964a1

fix fill_constant kernel bug test=develop (#2376) · f607870c

由 Wilber 提交于 11月 06, 2019

fill_constant kernel only registered float type, only the float data type is produced, which is obviously a bug.

Now, produce data based on the data type attr.

By the way, fix the cast kernel bug.

f607870c

H

change arm conv_kernels into basic to support mobilenetv1 test=develop (#2375) · 62bbf501
由 huzhiqiang 提交于 11月 06, 2019

62bbf501

05 11月, 2019 1 次提交
- L
  fix StepRNN model run related bugs (#2300) · 99d4f70e
  由 lijianshe02 提交于 11月 05, 2019
```
* fix step rnn model run bugs test=develop
```
  99d4f70e
04 11月, 2019 1 次提交
- H
  Move new op kernel into extra (#2348) · a438d0dc
  由 huzhiqiang 提交于 11月 04, 2019
```
* move some basic ops into extra type to reduce library size test=develop (#2347)
```
  a438d0dc
01 11月, 2019 2 次提交
- Z
  [XPU] add batchnorm op bridge and unit test (#2323) · b65861ee
  由 zhupengyang 提交于 11月 01, 2019
```
* [XPU] add batchnorm op bridge and unit test

test=develop

* fix DMLC_USE_GLOG

test=develop
```
  b65861ee
- T
  
  fix: fix conv_direct && test=develop (#2314) · ecca7325
  由 TianXiaogang 提交于 11月 01, 2019
  
  ecca7325
30 10月, 2019 2 次提交
- Z
  [XPU] update elemetwise_add, conv and mul ops (#2293) · 2a1856a9
  由 zhupengyang 提交于 10月 30, 2019
```
test=develop
```
  2a1856a9
- H
  
  [LITE][NPU] Use FullConnection op to solve the compatibility between Kirin 810 and 990 (#2283) · a5fe6f8e
  由 hong19860320 提交于 10月 30, 2019
  
  a5fe6f8e
29 10月, 2019 5 次提交
- L
  Add tanh op and gelu op for x86 platform (#2265) · e3368aa4
  由 liu zhengxi 提交于 10月 29, 2019
```
* add tanh op in x86 platform and its unittest, test=develop

* add gelu op on x86 platform and add its unittests, test=develop

* update depends for math_function for activation for gelu, test=develop
```
  e3368aa4
- Z
  [XPU] fix elementwise op bridge when x or y are from weight (#2272) · 40baeeae
  由 zhupengyang 提交于 10月 29, 2019
```
test=develop
```
  40baeeae
- H
  [LITE][NPU] Add supporting for Huawei offical DDK (#2262) · 92f0b3c5
  由 hong19860320 提交于 10月 29, 2019
```
* Add supporting for Huawei offical DDK
* Fix the param of graph op in NPU graph computing kernel
```
  92f0b3c5
- G
  modify layer_norm_compute_test.cc (#2261) · f2b9ca34
  由 guofei 提交于 10月 29, 2019
```
test=develop
```
  f2b9ca34
- Z
  [XPU] add mul op bridge (#2267) · 941c487d
  由 zhupengyang 提交于 10月 29, 2019
```
* [XPU] add mul op bridge and unit test

test=develop

* use tmp tensor for transposed y

test=develop
```
  941c487d
28 10月, 2019 2 次提交

Z
[XPU] add elementwise, pool, softmax op bridges and unit tests (#2264) · 231a658b
由 zhupengyang 提交于 10月 28, 2019
```
test=develop
```
231a658b

[LITE][XPU] initial support for XPU (#2202) · ac1b2f9f

由 hong19860320 提交于 10月 28, 2019

* Initial support for XPU
* Fix compiling errors of XPU
* Move XPU op kernel bridges from backends to kernels to fix deps order
* Change the namespace and directory of XPU bridges
* Add XPU SDK
* Fix header files and namespace of XPU SDK
* Add unit tests for relu and conv2d ops
* Restore the modification of paddle_api_test
* Supports simple model which contains only a relu layer
* Add compiling scripts for XPU
* Fix compiling errors of XPU
* Add comments for XPU LoadModel and BuildModel

ac1b2f9f

24 10月, 2019 1 次提交

Make inceptionv4, resnet50, googlenet can run on x86 paltform (#2250) · ca7fefa1

由 liu zhengxi 提交于 10月 24, 2019

* make inceptionv4, resnet50, googlenet can run on x86 paltform and fix the compare part in x86 unittests, test=develop

* fix googlenet tests for benchmark record, test=develop

* [framework][profile] fix profile dump bug when op is feed and fetch test=develop (sangoly)

ca7fefa1

23 10月, 2019 5 次提交
- W
  modify yolobox_cuda to support multiple runs (#2245) · e4b113eb
  由 Wilber 提交于 10月 23, 2019
```
* modify yolobox_cuda to support multiple runs test=develop
```
  e4b113eb
- T
  fix: fix fc batch bug (#2244) · ea4a5854
  由 TianXiaogang 提交于 10月 23, 2019
```
* fix: fix fc batch bug

* fix:fix fc bug;test=develop
```
  ea4a5854
- 石
  Add kernel version table and update framework.proto, test=develop (#2243) · 842326f7
  由石晓伟提交于 10月 23, 2019
```
* update framework.proto

* add compatibility check, test=develop

* remove head files, test=develop
```
  842326f7
- L
  Enable pool2d, dropout, transpose and transpose2 op on x86 (#2226) · 8006f5e6
  由 liu zhengxi 提交于 10月 23, 2019
```
* enable pool2d op on x86 and add its unit tests, test=develop

* enable dropout op and add its unit tests, test=develop

* add tranpose, transpose2 op and add their unit tests, test=develop
```
  8006f5e6
- S
  add python api (#2225) · bbfaacec
  由 sangoly 提交于 10月 23, 2019
```
* [python api] init add python api test=develop
```
  bbfaacec
22 10月, 2019 1 次提交

Transformer pr (#2214) · 330644b0

由 TianXiaogang 提交于 10月 22, 2019

* feat: add beam_search_special function for support nlp model

* fix: add beam_search_compute kernel input and output

* feat: add assign op & copy_compute kernel

* feat: add fill_const_batch_size_like op & kernel

* feat: add layer_norm op and kernel and ut

* fix: fix some bugs
    fix mul_op infer_shape bug when x_dim_idx = 2, x_dims.size()=3 & y_dim_idx = 1, y_dims.size()=2
    fix elementwise_compute bug when y axis is all 1
    fix beam_search choose math_func wrong bug
    fix layer_norm get attr bug
    fix fill_constant_batch_size_like shape_set bug

* feat: add gather op and kernel & and transform ut

* feats: add ops and fix bugs to support transformer op
       fix type_cast passes to skip `while`
       fix elementwise infer_shape bug when x.dims=3 and y.dims={1} & axis=0
       fix lookup_table compute bug
       fix read_from_array/beam_search/increment/compate/gather ops data_type problems

* fix:
    transfomer ut add word read inferface
    fix copy/gather/norm/layer_norm include path problem

* fix:debug info

* fix: fix input reshape bug

* fix: fix norm bug

* style: style fix & test=develop

* style: fix operators cmakelist

* style: fix operators cmakelist; test=develop

* fix and test=develop

* fix and test=develop

* style: style fix; test=develop

330644b0

21 10月, 2019 2 次提交

to support yolov3 unet alexnet can run on tx2 (#2216) · d3d6ed4b
由 myq406450149 提交于 10月 21, 2019
```
* add gpu kernel mul pool relu scale softmax dropout bilinear_interp and can run in tx2

* rm GREATER_EQUAL
```
d3d6ed4b

add cuda op(pool & softmax), support conv with padding_algorithm · 1a18d682

由 yiicy 提交于 10月 21, 2019

* cuda add softmax and pool op

* * fix armlinux can find sys/system_properties.h
* conv add padding_algorithm
test=develop

* delete padding_algorithm in op param, test=develop

* fix bugs, test=develop

1a18d682

18 10月, 2019 1 次提交
- W
  fix yolobox_cuda_test (#2208) · 2f57f5b4
  由 Wilber 提交于 10月 18, 2019
```
fix yolobox_cuda test precision error
```
  2f57f5b4
17 10月, 2019 4 次提交

J

add bilinear_interp_cuda_op, test=develop (#2197) · cb6b1b1c
由 juncaipeng 提交于 10月 17, 2019

cb6b1b1c

fix npu path (#2210) · c10a8b16

由 zhupengyang 提交于 10月 17, 2019

* move lite/backends/npu/bridges --> lite/kernels/npu/

test=develop

* fix namespace for npu

test=develop

* mv npu runtime file to lite/backends/npu

test=develop

c10a8b16

speedup fp32 depthwise conv · c5cd78ab

由 HappyAngel 提交于 10月 17, 2019

* update con_dw

* update

* add conv_depthwise_3x3s1.cc and conv_depthwise_3x3s2.cc

* add conv_depthwise_3x3s1_fp32 and conv_depthwise_3x3s2_fp32

* add new conv_dw

* only support conv_dw pad=0, 1

* add conv_dw_s1 conv_dw_s2 fp32

*     //conv2_func _impl2{nullptr};
update conv_dw, add conv_3x3s1 and conv_3x3s2, pad=[0,1]

* fix format, test=develop

* fix formmat, test=develop

c5cd78ab

L

enable batch_norm op and add its unit tests, test=develop (#2201) · 95372548
由 liu zhengxi 提交于 10月 17, 2019

95372548

16 10月, 2019 1 次提交
- L
  enable conv2d op and its unit tests, test=develop (#2200) · 0fed350e
  由 liu zhengxi 提交于 10月 16, 2019
```
enable conv2d op and its unit tests on x86 device
```
  0fed350e