提交 · ac0f4500e793f260f6100c12f3e207f92466f501 · PaddlePaddle / Paddle-Lite

07 11月, 2019 1 次提交

Cherry pick fix_arm_kernel type (#2389) · ac0f4500

由 huzhiqiang 提交于 11月 06, 2019

* change arm conv_kernels into basic to support mobilenetv1 test=develop (#2375)
* check arm kernels type to make sure all_library_links work normally (#2386)

ac0f4500

01 11月, 2019 2 次提交
- H
  
  move some basic ops into extra type to reduce library size test=develop (#2347) · 30f17891
  由 huzhiqiang 提交于 11月 01, 2019
  
  30f17891
- T
  fix: fix conv_direct && test=develop (#2314) (#2324) · c18b6ebc
  由 TianXiaogang 提交于 11月 01, 2019
```
cherry-pick
```
  c18b6ebc
29 10月, 2019 1 次提交
- G
  modify layer_norm_compute_test.cc (#2261) (#2277) · 850952e9
  由 guofei 提交于 10月 29, 2019
```
test=develop
```
  850952e9
24 10月, 2019 1 次提交

Make inceptionv4, resnet50, googlenet can run on x86 paltform (#2250) · edb4ea9a

由 liu zhengxi 提交于 10月 24, 2019

* make inceptionv4, resnet50, googlenet can run on x86 paltform and fix the compare part in x86 unittests, test=develop

* fix googlenet tests for benchmark record, test=develop

* [framework][profile] fix profile dump bug when op is feed and fetch test=develop (sangoly)

edb4ea9a

23 10月, 2019 5 次提交
- W
  modify yolobox_cuda to support multiple runs (#2245) · a4a19ba4
  由 Wilber 提交于 10月 23, 2019
```
* modify yolobox_cuda to support multiple runs test=develop
```
  a4a19ba4
- T
  fix: fix fc batch bug (#2244) · b2db5f49
  由 TianXiaogang 提交于 10月 23, 2019
```
* fix: fix fc batch bug

* fix:fix fc bug;test=develop
```
  b2db5f49
- 石
  Add kernel version table and update framework.proto, test=develop (#2243) · 362275ed
  由石晓伟提交于 10月 23, 2019
```
* update framework.proto

* add compatibility check, test=develop

* remove head files, test=develop
```
  362275ed
- L
  Enable pool2d, dropout, transpose and transpose2 op on x86 (#2226) · aa507f9b
  由 liu zhengxi 提交于 10月 23, 2019
```
* enable pool2d op on x86 and add its unit tests, test=develop

* enable dropout op and add its unit tests, test=develop

* add tranpose, transpose2 op and add their unit tests, test=develop
```
  aa507f9b
- S
  add python api (#2225) · 76e74ef1
  由 sangoly 提交于 10月 23, 2019
```
* [python api] init add python api test=develop
```
  76e74ef1
22 10月, 2019 1 次提交

Transformer pr (#2214) · f0a6c1eb

由 TianXiaogang 提交于 10月 22, 2019

* feat: add beam_search_special function for support nlp model

* fix: add beam_search_compute kernel input and output

* feat: add assign op & copy_compute kernel

* feat: add fill_const_batch_size_like op & kernel

* feat: add layer_norm op and kernel and ut

* fix: fix some bugs
    fix mul_op infer_shape bug when x_dim_idx = 2, x_dims.size()=3 & y_dim_idx = 1, y_dims.size()=2
    fix elementwise_compute bug when y axis is all 1
    fix beam_search choose math_func wrong bug
    fix layer_norm get attr bug
    fix fill_constant_batch_size_like shape_set bug

* feat: add gather op and kernel & and transform ut

* feats: add ops and fix bugs to support transformer op
       fix type_cast passes to skip `while`
       fix elementwise infer_shape bug when x.dims=3 and y.dims={1} & axis=0
       fix lookup_table compute bug
       fix read_from_array/beam_search/increment/compate/gather ops data_type problems

* fix:
    transfomer ut add word read inferface
    fix copy/gather/norm/layer_norm include path problem

* fix:debug info

* fix: fix input reshape bug

* fix: fix norm bug

* style: style fix & test=develop

* style: fix operators cmakelist

* style: fix operators cmakelist; test=develop

* fix and test=develop

* fix and test=develop

* style: style fix; test=develop

f0a6c1eb

21 10月, 2019 2 次提交

to support yolov3 unet alexnet can run on tx2 (#2216) · 57d8e42e
由 myq406450149 提交于 10月 21, 2019
```
* add gpu kernel mul pool relu scale softmax dropout bilinear_interp and can run in tx2

* rm GREATER_EQUAL
```
57d8e42e

add cuda op(pool & softmax), support conv with padding_algorithm · 305130fc

由 yiicy 提交于 10月 21, 2019

* cuda add softmax and pool op

* * fix armlinux can find sys/system_properties.h
* conv add padding_algorithm
test=develop

* delete padding_algorithm in op param, test=develop

* fix bugs, test=develop

305130fc

18 10月, 2019 1 次提交
- W
  fix yolobox_cuda_test (#2208) · 2a6a259d
  由 Wilber 提交于 10月 18, 2019
```
fix yolobox_cuda test precision error
```
  2a6a259d
17 10月, 2019 4 次提交

J

add bilinear_interp_cuda_op, test=develop (#2197) · 4ac51a6b
由 juncaipeng 提交于 10月 17, 2019

4ac51a6b

fix npu path (#2210) · 7c722a37

由 zhupengyang 提交于 10月 17, 2019

* move lite/backends/npu/bridges --> lite/kernels/npu/

test=develop

* fix namespace for npu

test=develop

* mv npu runtime file to lite/backends/npu

test=develop

7c722a37

speedup fp32 depthwise conv · 2f6d5f9e

由 HappyAngel 提交于 10月 17, 2019

* update con_dw

* update

* add conv_depthwise_3x3s1.cc and conv_depthwise_3x3s2.cc

* add conv_depthwise_3x3s1_fp32 and conv_depthwise_3x3s2_fp32

* add new conv_dw

* only support conv_dw pad=0, 1

* add conv_dw_s1 conv_dw_s2 fp32

*     //conv2_func _impl2{nullptr};
update conv_dw, add conv_3x3s1 and conv_3x3s2, pad=[0,1]

* fix format, test=develop

* fix formmat, test=develop

2f6d5f9e

L

enable batch_norm op and add its unit tests, test=develop (#2201) · a3241ca7
由 liu zhengxi 提交于 10月 17, 2019

a3241ca7

16 10月, 2019 1 次提交
- L
  enable conv2d op and its unit tests, test=develop (#2200) · 459848c4
  由 liu zhengxi 提交于 10月 16, 2019
```
enable conv2d op and its unit tests on x86 device
```
  459848c4
15 10月, 2019 2 次提交

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of... · 72c11758

由 Yuan Shuai 提交于 10月 15, 2019

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of CONVERT_TYPE_TO and READ/WRITE image, memory reuse in ResetLazyImage2D (#2170)

* add macro of CONVERT_TYPE_TO and READ/WRITE image. test=develop

* add data type control. test=develop

* fix io op as general layout and precision. test=develop

* Fix memory reuse strategy for opencl image2d. test=develop

* remove std::array, std::map in about opencl backend. test=develop

72c11758

[NPU] Fix and refine the supporting of multi NPU models (#2037) · 7a731b7f

由 hong19860320 提交于 10月 15, 2019

* [NPU] Fix the bug of loading multi NPU models
test=develop

* [NPU] Use lite tensor to store NPU model, fix the management of multi NPU models, support loading NPU model from memory and reduce the modification of framework
test=develop

* [NPU] Remove redundant header files for NPU bridges,
test=develop

* [NPU] fix NPU deps
test=develop

* [NPU] refine the compiling script for NPU
test=develop

* [NPU] remove redundant subdirectory in lite/CMakeLists.txt
test=develop

* [NPU] Fix and refine NPU test case
test=develop

* [NPU] revoke the modification of other non-NPU modules
test=develop

* [NPU] Remove NPU bridges if target is tiny publish
test=develop

7a731b7f

14 10月, 2019 3 次提交
- J
  fix bug for reshape op, test=develop (#2141) · 421c6305
  由 juncaipeng 提交于 10月 14, 2019
```
* fix bug for reshape op, test=develop
```
  421c6305
- Z
  align yolov3 cuda int8 (#2183) · 80d35725
  由 Zhaolong Xing 提交于 10月 14, 2019
```
test=develop
```
  80d35725
- L
  fix asr modle related kernel bugs test=develop (#2179) · 792d898a
  由 lijianshe02 提交于 10月 14, 2019
```
* fix asr modle related kernel bugs test=develop
```
  792d898a
12 10月, 2019 1 次提交

fix conv_transpose error (#2165) · a6b1e4fa

由 Xiaoyang LI 提交于 10月 12, 2019

* fix conv_transpose error

* fix build error, enable basic test of conv_transpose, test=develop

a6b1e4fa

11 10月, 2019 3 次提交

J

add rsqrt op, test=develop (#2176) · dfce4621
由 juncaipeng 提交于 10月 11, 2019

dfce4621

CUDA: can run yolov3 int8 (#2172) · 7931104f

由 Zhaolong Xing 提交于 10月 11, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

* can run yolov3 int8 test=develop

7931104f

[LITE][OPENCL] support image2d type (#2158) · 77cdbdce

由 Yuan Shuai 提交于 10月 11, 2019

* [LITE][OPENCL] support image2d. test=develop

* add context changed with consider image*. test=develop

* add layout, relu image kernels. test=develop

* replace image_data with data, mutable_image_data with mutable_data, test=develop

* comment unused var. test=develop

* remove unused var. test=develop

77cdbdce

10 10月, 2019 1 次提交
- W
  fix yolobox_cuda bug · f4ac2768
  由 Wilber 提交于 10月 10, 2019
```
* fix yolobox_cuda bug 
* update code format
```
  f4ac2768
09 10月, 2019 1 次提交

improve dw conv performance · 4b9df8fb

由 yiicy 提交于 10月 09, 2019

*  imporve prepack_input func speed in int8 3x3s1 dw conv

* fix code style

* fix code style

* improve 3x3s1 dw fp32 conv speed a little

* arm add 5x5s1 int8 dw conv, test=develop

4b9df8fb

27 9月, 2019 1 次提交

can run yolov3 fp32 on cuda devices (#2092) · 3d6d744f

由 Zhaolong Xing 提交于 9月 27, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

3d6d744f

25 9月, 2019 1 次提交
- X
  
  add workspace compute funcs for direct conv, test=develop (#2132) · 61647c35
  由 Xiaoyang LI 提交于 9月 25, 2019
  
  61647c35
23 9月, 2019 1 次提交
- J
  add cast from uint8 to float, test=develop (#2080) · 9941d746
  由 juncaipeng 提交于 9月 23, 2019
```
* add cast from uint8 to float, test=develop
```
  9941d746
20 9月, 2019 1 次提交
- P
  
  refine concat cuda kernel, test=develop (#2081) · cef884e5
  由 Pei Yang 提交于 9月 20, 2019
  
  cef884e5
19 9月, 2019 1 次提交

石

add full_api_static target and fix building errors, test=develop (#2064) · eef7ea0f

由石晓伟提交于 9月 19, 2019

* add full_api_static target and fix building errors, test=develop

* fix build errors, test=develop

* fix code style, test=develop

* fix lite/model_parser/pb/var_desc.cc, test=develop

* fix building errors, test=develop

* modify lite/tools/debug/CMakeLists.txt, test=develop

eef7ea0f

18 9月, 2019 1 次提交

fix bias quantize error && fix clang build error (#2049) · 81dffbe8

由 Xiaoyang LI 提交于 9月 18, 2019

* fix gemm_int8, gemv-int8 and conv-int8 math function, add float bias

* change conv impl

* neon int8 kernel support float bias

* arm compute kernel support float bias

* add math_test target

* add tensor utils for testing, fix sgemm ut error

* add gemm_int8 unit test, support float bias

* fix build script

* add conv compute unit test for arm

* fix build script, test=develop

* fix fp32 dw conv3x3s1, test=develop

* add fp32 dw conv3x3s1, test=develop

* add armv7 fp32 dw conv3x3s1, test=develop

* add fp32 depthwise conv3x3s2, test=develop

* fix fp32 conv3x3 depthwise build error, test=develop

* fix gemm_like conv trans weights error, test=develop

* fix int8 depthwise conv3x3 error, test=develop

* turn on all test for arm fp32 conv, test=develop

* fix int8 conv1x1 error

* fix int8 direct conv3x3s1 error, test=develop

* fix int8 direct conv3x3s2, test=develop

* turn on all test for arm int8 conv, test=develop

* fix int8 fc error, change mobilenetv1-int8 ground-truth result to fluid, test=develop

* remove debug info, strip ut binary, test=develop

* fix conv compute error, test=develop

* change Init() to ReInitWhenNeeded(), test=develop

* fix code style, test=develop

* remote engine_test, test=develop

* fix building server tests error, test=develop

* fix sdot clang build error, test=develop

* fix sgemm ut timeout error, test=develop

* fix clang build error, test=develop

* turn off math basic test due to ci time out, test=develop

* fix conv_int8 ut error, test=develop

81dffbe8

17 9月, 2019 2 次提交
- W
  modify norm kernel to run caffe_facedetection model (#2008) · 71bb3188
  由 Wilber 提交于 9月 17, 2019
```
* modify norm kernel to run caffe_facedetection model

* reserve bind norm, remove calc of norm output
```
  71bb3188
- L
  add fill_constant_batch_size_like op and add its unittest (#2044) · f5338469
  由 liu zhengxi 提交于 9月 17, 2019
```
* add fill_constant_batch_size_like op and add its unittest
```
  f5338469
16 9月, 2019 1 次提交
- L
  Gru op (#2002) · eb42f9ee
  由 lhl960107 提交于 9月 16, 2019
```
* add x86 gru&&relu&&sequence_expand_as op test=develop
```
  eb42f9ee
12 9月, 2019 1 次提交
- L
  
  add matmul op kernels for asr test=develop (#2032) · b6379fed
  由 lijianshe02 提交于 9月 12, 2019
  
  b6379fed