提交 · 31f1e3827775f5823e788a1cf7acc8f03e35ed39 · PaddlePaddle / Paddle-Lite

21 10月, 2019 6 次提交
- 石
  link static library with cuda, test=develop (#2228) · 31f1e382
  由石晓伟提交于 10月 21, 2019
```
* add static libraries of cuda, test=develop

* update cuda make
```
  31f1e382
- to support yolov3 unet alexnet can run on tx2 (#2216) · d3d6ed4b
  由 myq406450149 提交于 10月 21, 2019
```
* add gpu kernel mul pool relu scale softmax dropout bilinear_interp and can run in tx2

* rm GREATER_EQUAL
```
  d3d6ed4b
- J
  
  add GetExceptionMsg for paddle_inference_api. test=develop (#2231) · f623b1d4
  由 Jiaying Zhao 提交于 10月 21, 2019
  
  f623b1d4
- Y
  add cuda op(pool & softmax), support conv with padding_algorithm · 1a18d682
  由 yiicy 提交于 10月 21, 2019
```
* cuda add softmax and pool op

* * fix armlinux can find sys/system_properties.h
* conv add padding_algorithm
test=develop

* delete padding_algorithm in op param, test=develop

* fix bugs, test=develop
```
  1a18d682
- X
  
  fix ios build error, test=develop (#2222) · aefe9323
  由 Xiaoyang LI 提交于 10月 21, 2019
  
  aefe9323
- H
  Fix ‘Large memory usage of Naive model loading’ (#2175) · b0f94214
  由 huzhiqiang 提交于 10月 21, 2019
```
Fix ‘Large memory usage of Naive model loading’  (#2175)
```
  b0f94214
18 10月, 2019 2 次提交

W
fix yolobox_cuda_test (#2208) · 2f57f5b4
由 Wilber 提交于 10月 18, 2019
```
fix yolobox_cuda test precision error
```
2f57f5b4

Fix codestyle of GetInputName&GetOutputName (#2185) · 27a40b8f

由 huzhiqiang 提交于 10月 18, 2019

* add shell file to automatically build and collect publish result test=develop

* modify codestyle of getInputNames test=develop

* test=develop

* rm publish.sh

* remove copy of func param

* test=develop

* test=devcelop

* test=develop

* test=develop

* const & test=develop

* modify variable defination test=develop

* test=develop

* test=develop

* test=develop

* test=develop

27a40b8f

17 10月, 2019 5 次提交

J

add bilinear_interp_cuda_op, test=develop (#2197) · cb6b1b1c
由 juncaipeng 提交于 10月 17, 2019

cb6b1b1c
S

fix “CL_INVALID_KERNEL_ARGS ” error， test=develop (#2213) · 0baf7c05
由 StarryRain 提交于 10月 17, 2019

0baf7c05

fix npu path (#2210) · c10a8b16

由 zhupengyang 提交于 10月 17, 2019

* move lite/backends/npu/bridges --> lite/kernels/npu/

test=develop

* fix namespace for npu

test=develop

* mv npu runtime file to lite/backends/npu

test=develop

c10a8b16

speedup fp32 depthwise conv · c5cd78ab

由 HappyAngel 提交于 10月 17, 2019

* update con_dw

* update

* add conv_depthwise_3x3s1.cc and conv_depthwise_3x3s2.cc

* add conv_depthwise_3x3s1_fp32 and conv_depthwise_3x3s2_fp32

* add new conv_dw

* only support conv_dw pad=0, 1

* add conv_dw_s1 conv_dw_s2 fp32

*     //conv2_func _impl2{nullptr};
update conv_dw, add conv_3x3s1 and conv_3x3s2, pad=[0,1]

* fix format, test=develop

* fix formmat, test=develop

c5cd78ab

L

enable batch_norm op and add its unit tests, test=develop (#2201) · 95372548
由 liu zhengxi 提交于 10月 17, 2019

95372548

16 10月, 2019 5 次提交
- Z
  Ban feed and fetch op during inference (#2198) · ad541652
  由 Zhaolong Xing 提交于 10月 16, 2019
```
* init: delete feed and fetch op, using zero copy
test=develop

* delete the unused test
test=develop
```
  ad541652
- J
  
  Open merge_cl_to_so switch and delete -I(cl_path) build option. test=develop (#2206) · eb7c5829
  由 Jiaying Zhao 提交于 10月 16, 2019
  
  eb7c5829
- L
  enable conv2d op and its unit tests, test=develop (#2200) · 0fed350e
  由 liu zhengxi 提交于 10月 16, 2019
```
enable conv2d op and its unit tests on x86 device
```
  0fed350e
- X
  
  support global pooling ... test=develop (#2204) · 9f867daf
  由 xiebaiyuan 提交于 10月 16, 2019
  
  9f867daf
- S
  [framework][place] remove prefered_place and kHost in valid_places (#2192) · 17833acb
  由 sangoly 提交于 10月 16, 2019
```
* [framework][place] remove prefered_place, use place order in valid_place array instead test=develop

* remove kHost from valid_places test=develop
```
  17833acb
15 10月, 2019 6 次提交

J
Fix quant dequant fuse pass (#2190) · 31ab471e
由 juncaipeng 提交于 10月 15, 2019
```
* fix bug for accessing the removed node, test=develop
```
31ab471e
J
fix benchmark, test=develop (#2188) · dbb660ee
由 juncaipeng 提交于 10月 15, 2019
```
* fix benchmark, test=develop
```
dbb660ee
Y

fix persistable test=develop (#2191) · 92d0552f
由 Yanzhan Yang 提交于 10月 15, 2019

92d0552f
石

fix pass selection, test=develop (#2187) · a1a22a69
由石晓伟提交于 10月 15, 2019

a1a22a69

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of... · 26d05370

由 Yuan Shuai 提交于 10月 15, 2019

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of CONVERT_TYPE_TO and READ/WRITE image, memory reuse in ResetLazyImage2D (#2170)

* add macro of CONVERT_TYPE_TO and READ/WRITE image. test=develop

* add data type control. test=develop

* fix io op as general layout and precision. test=develop

* Fix memory reuse strategy for opencl image2d. test=develop

* remove std::array, std::map in about opencl backend. test=develop

26d05370

[NPU] Fix and refine the supporting of multi NPU models (#2037) · e184d474

由 hong19860320 提交于 10月 15, 2019

* [NPU] Fix the bug of loading multi NPU models
test=develop

* [NPU] Use lite tensor to store NPU model, fix the management of multi NPU models, support loading NPU model from memory and reduce the modification of framework
test=develop

* [NPU] Remove redundant header files for NPU bridges,
test=develop

* [NPU] fix NPU deps
test=develop

* [NPU] refine the compiling script for NPU
test=develop

* [NPU] remove redundant subdirectory in lite/CMakeLists.txt
test=develop

* [NPU] Fix and refine NPU test case
test=develop

* [NPU] revoke the modification of other non-NPU modules
test=develop

* [NPU] Remove NPU bridges if target is tiny publish
test=develop

e184d474

14 10月, 2019 5 次提交
- J
  fix bug for reshape op, test=develop (#2141) · 4a061a27
  由 juncaipeng 提交于 10月 14, 2019
```
* fix bug for reshape op, test=develop
```
  4a061a27
- Z
  align yolov3 cuda int8 (#2183) · ed38d79b
  由 Zhaolong Xing 提交于 10月 14, 2019
```
test=develop
```
  ed38d79b
- H
  add GetInputNames 、 GetOutPutNames 、 GetInputByName and GetTensor method (#2154) · 1cd077dc
  由 huzhiqiang 提交于 10月 14, 2019
```
* add GetInputNames and GetOutPutNames and GetInputByName method test=develop
```
  1cd077dc
- L
  fix asr modle related kernel bugs test=develop (#2179) · 2f035fec
  由 lijianshe02 提交于 10月 14, 2019
```
* fix asr modle related kernel bugs test=develop
```
  2f035fec
- J
  Optimize quant_dequant_fuse_pass (#2169) · 0260d322
  由 juncaipeng 提交于 10月 14, 2019
```
* optimize quant_dequant_fuse_pass, test=develop
```
  0260d322
12 10月, 2019 2 次提交
- J
  
  fix clang compile error. test=develop (#2180) · 9aa795ca
  由 Jiaying Zhao 提交于 10月 12, 2019
  
  9aa795ca
- X
  fix conv_transpose error (#2165) · 9a464d63
  由 Xiaoyang LI 提交于 10月 12, 2019
```
* fix conv_transpose error

* fix build error, enable basic test of conv_transpose, test=develop
```
  9a464d63
11 10月, 2019 5 次提交

J

add rsqrt op, test=develop (#2176) · 78ddd64d
由 juncaipeng 提交于 10月 11, 2019

78ddd64d
Y

1. fix group logic for convolution op. 2. add pixel shuffle op for OpenCL. (#2178) · e0aafc03
由 Yanzhan Yang 提交于 10月 11, 2019

e0aafc03

CUDA: can run yolov3 int8 (#2172) · 29f448c6

由 Zhaolong Xing 提交于 10月 11, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

* can run yolov3 int8 test=develop

29f448c6

move the method of SetThread and SetPowerMode from MobileConfig into ConfigBase (#2147) · 6aba3b8b

由 huzhiqiang 提交于 10月 11, 2019

* move the method of SetThread and SetPowerMode from MobileConfig into ConfigBase 
* cxxPredictor will support SetThread and SetPowerMode method

6aba3b8b

[LITE][OPENCL] support image2d type (#2158) · 74074146

由 Yuan Shuai 提交于 10月 11, 2019

* [LITE][OPENCL] support image2d. test=develop

* add context changed with consider image*. test=develop

* add layout, relu image kernels. test=develop

* replace image_data with data, mutable_image_data with mutable_data, test=develop

* comment unused var. test=develop

* remove unused var. test=develop

74074146

10 10月, 2019 4 次提交
- W
  fix yolobox_cuda bug · 8bc7c043
  由 Wilber 提交于 10月 10, 2019
```
* fix yolobox_cuda bug 
* update code format
```
  8bc7c043
- Y
  1. improve n-fold quantification algorithm by introducing a minimal size for... · 148ef22a
  由 Yanzhan Yang 提交于 10月 10, 2019
```
1. improve n-fold quantification algorithm by introducing a minimal size for each fold. 2. automatically search for best n for n-fold algorithm. (#2167)
```
  148ef22a
- Y
  [CMAKE] Abandon strip when CMAKE_BUILD_TYPE=Debug (#2155) · 9f73c064
  由 Yuan Shuai 提交于 10月 10, 2019
```
* [CMAKE] Abandon strip when CMAKE_BUILD_TYPE=Debug

* Add CMAKE_BUILD_TYPE info in cmake. test=develop
```
  9f73c064
- X
  
  fix an calc bug in test-mobilenetgpu (#2162) · 5245133a
  由 xiebaiyuan 提交于 10月 10, 2019
  
  5245133a