提交 · 2f6d5f9e5ef9cd9ba02114855273dee40ac1774d · PaddlePaddle / Paddle-Lite

17 10月, 2019 1 次提交

由 HappyAngel 提交于 10月 17, 2019

* update con_dw

* update

* add conv_depthwise_3x3s1.cc and conv_depthwise_3x3s2.cc

* add conv_depthwise_3x3s1_fp32 and conv_depthwise_3x3s2_fp32

* add new conv_dw

* only support conv_dw pad=0, 1

* add conv_dw_s1 conv_dw_s2 fp32

*     //conv2_func _impl2{nullptr};
update conv_dw, add conv_3x3s1 and conv_3x3s2, pad=[0,1]

* fix format, test=develop

* fix formmat, test=develop

2f6d5f9e

16 10月, 2019 1 次提交
- Z
  Ban feed and fetch op during inference (#2198) · 75e8a6fc
  由 Zhaolong Xing 提交于 10月 16, 2019
```
* init: delete feed and fetch op, using zero copy
test=develop

* delete the unused test
test=develop
```
  75e8a6fc
15 10月, 2019 2 次提交

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of... · 72c11758

由 Yuan Shuai 提交于 10月 15, 2019

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of CONVERT_TYPE_TO and READ/WRITE image, memory reuse in ResetLazyImage2D (#2170)

* add macro of CONVERT_TYPE_TO and READ/WRITE image. test=develop

* add data type control. test=develop

* fix io op as general layout and precision. test=develop

* Fix memory reuse strategy for opencl image2d. test=develop

* remove std::array, std::map in about opencl backend. test=develop

72c11758

[NPU] Fix and refine the supporting of multi NPU models (#2037) · 7a731b7f

由 hong19860320 提交于 10月 15, 2019

* [NPU] Fix the bug of loading multi NPU models
test=develop

* [NPU] Use lite tensor to store NPU model, fix the management of multi NPU models, support loading NPU model from memory and reduce the modification of framework
test=develop

* [NPU] Remove redundant header files for NPU bridges,
test=develop

* [NPU] fix NPU deps
test=develop

* [NPU] refine the compiling script for NPU
test=develop

* [NPU] remove redundant subdirectory in lite/CMakeLists.txt
test=develop

* [NPU] Fix and refine NPU test case
test=develop

* [NPU] revoke the modification of other non-NPU modules
test=develop

* [NPU] Remove NPU bridges if target is tiny publish
test=develop

7a731b7f

14 10月, 2019 2 次提交
- Z
  align yolov3 cuda int8 (#2183) · 80d35725
  由 Zhaolong Xing 提交于 10月 14, 2019
```
test=develop
```
  80d35725
- L
  fix asr modle related kernel bugs test=develop (#2179) · 792d898a
  由 lijianshe02 提交于 10月 14, 2019
```
* fix asr modle related kernel bugs test=develop
```
  792d898a
11 10月, 2019 3 次提交

J

add rsqrt op, test=develop (#2176) · dfce4621
由 juncaipeng 提交于 10月 11, 2019

dfce4621

CUDA: can run yolov3 int8 (#2172) · 7931104f

由 Zhaolong Xing 提交于 10月 11, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

* can run yolov3 int8 test=develop

7931104f

[LITE][OPENCL] support image2d type (#2158) · 77cdbdce

由 Yuan Shuai 提交于 10月 11, 2019

* [LITE][OPENCL] support image2d. test=develop

* add context changed with consider image*. test=develop

* add layout, relu image kernels. test=develop

* replace image_data with data, mutable_image_data with mutable_data, test=develop

* comment unused var. test=develop

* remove unused var. test=develop

77cdbdce

10 10月, 2019 1 次提交
- W
  fix yolobox_cuda bug · f4ac2768
  由 Wilber 提交于 10月 10, 2019
```
* fix yolobox_cuda bug 
* update code format
```
  f4ac2768
09 10月, 2019 1 次提交

improve dw conv performance · 4b9df8fb

由 yiicy 提交于 10月 09, 2019

*  imporve prepack_input func speed in int8 3x3s1 dw conv

* fix code style

* fix code style

* improve 3x3s1 dw fp32 conv speed a little

* arm add 5x5s1 int8 dw conv, test=develop

4b9df8fb

27 9月, 2019 1 次提交

can run yolov3 fp32 on cuda devices (#2092) · 3d6d744f

由 Zhaolong Xing 提交于 9月 27, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

3d6d744f

25 9月, 2019 1 次提交
- X
  
  add workspace compute funcs for direct conv, test=develop (#2132) · 61647c35
  由 Xiaoyang LI 提交于 9月 25, 2019
  
  61647c35
19 9月, 2019 3 次提交

石

add full_api_static target and fix building errors, test=develop (#2064) · eef7ea0f

由石晓伟提交于 9月 19, 2019

* add full_api_static target and fix building errors, test=develop

* fix build errors, test=develop

* fix code style, test=develop

* fix lite/model_parser/pb/var_desc.cc, test=develop

* fix building errors, test=develop

* modify lite/tools/debug/CMakeLists.txt, test=develop

eef7ea0f

fix building model_optimize_tool error on mac (#2075) · 80e4172e

由 Xiaoyang LI 提交于 9月 19, 2019

* fix building model_optimize_tool error on mac, test=develop

* fix model_optimize_tool build error, test=develop

80e4172e

Bug fix for model save and load (#1992) · 8efbdc66

由 TianXiaogang 提交于 9月 19, 2019

* fix: fix model parser and save bug

* style: delete debug code

* fix: fix light_predictor program run model with subblock bug

8efbdc66

18 9月, 2019 1 次提交

fix bias quantize error && fix clang build error (#2049) · 81dffbe8

由 Xiaoyang LI 提交于 9月 18, 2019

* fix gemm_int8, gemv-int8 and conv-int8 math function, add float bias

* change conv impl

* neon int8 kernel support float bias

* arm compute kernel support float bias

* add math_test target

* add tensor utils for testing, fix sgemm ut error

* add gemm_int8 unit test, support float bias

* fix build script

* add conv compute unit test for arm

* fix build script, test=develop

* fix fp32 dw conv3x3s1, test=develop

* add fp32 dw conv3x3s1, test=develop

* add armv7 fp32 dw conv3x3s1, test=develop

* add fp32 depthwise conv3x3s2, test=develop

* fix fp32 conv3x3 depthwise build error, test=develop

* fix gemm_like conv trans weights error, test=develop

* fix int8 depthwise conv3x3 error, test=develop

* turn on all test for arm fp32 conv, test=develop

* fix int8 conv1x1 error

* fix int8 direct conv3x3s1 error, test=develop

* fix int8 direct conv3x3s2, test=develop

* turn on all test for arm int8 conv, test=develop

* fix int8 fc error, change mobilenetv1-int8 ground-truth result to fluid, test=develop

* remove debug info, strip ut binary, test=develop

* fix conv compute error, test=develop

* change Init() to ReInitWhenNeeded(), test=develop

* fix code style, test=develop

* remote engine_test, test=develop

* fix building server tests error, test=develop

* fix sdot clang build error, test=develop

* fix sgemm ut timeout error, test=develop

* fix clang build error, test=develop

* turn off math basic test due to ci time out, test=develop

* fix conv_int8 ut error, test=develop

81dffbe8

17 9月, 2019 1 次提交
- J
  fix yolo_box bug (#2034) · 1b1d7a83
  由 juncaipeng 提交于 9月 17, 2019
```
* fix yolo_box bug, test=develop

* fix test bug for yolo_box, test=develop
```
  1b1d7a83
16 9月, 2019 2 次提交
- L
  Gru op (#2002) · eb42f9ee
  由 lhl960107 提交于 9月 16, 2019
```
* add x86 gru&&relu&&sequence_expand_as op test=develop
```
  eb42f9ee
- X
  
  fix math dependencies error (#2023) · 5900c784
  由 Xiaoyang LI 提交于 9月 16, 2019
  
  5900c784
12 9月, 2019 4 次提交
- Z
  fix bilinear-interp arm compute (#2029) · ca424e73
  由 zhupengyang 提交于 9月 12, 2019
```
fix bilinear-interp unit test for more cases

test=develop
```
  ca424e73
- H
  add x86 math lstm and selected_rows test=develop (#1991) · dc27e9ff
  由 huzhiqiang 提交于 9月 12, 2019
```
add math function: lstm  and selected_rows into lite/x86/math
add selected_rows and rw_lock into lite/fluid
add lstm_cpu_kernel and  lstm_kernel into lite/x86/detail
```
  dc27e9ff
- W
  
  add min_max_aspect_ratios_order attr in prior box op test=develop (#2016) · cf84d42b
  由 Wilber 提交于 9月 12, 2019
  
  cf84d42b
- W
  add transpose kernel for cuda test=develop (#1997) · cba5736f
  由 Wilber 提交于 9月 12, 2019
```
add transpose kernel for cuda
```
  cba5736f
11 9月, 2019 1 次提交
- 石
  make passes related to the device type, test=develop (#2012) · 8ca10db8
  由石晓伟提交于 9月 11, 2019
```
* make passes related to the device type, test=develop

* improve tips, test=develop
```
  8ca10db8
10 9月, 2019 3 次提交
- L
  
  add x86 softmax kernel and fix jit compute bugs test=develop (#2007) · 81132a32
  由 lijianshe02 提交于 9月 10, 2019
  
  81132a32
- W
  
  add elementwise_sub and modify argmax (#1964) · 62ea82d0
  由 Wilber 提交于 9月 10, 2019
  
  62ea82d0
- T
  
  fix fpga compile problem and kernels (#1989) · 0720653b
  由 TianXiaogang 提交于 9月 10, 2019
  
  0720653b
09 9月, 2019 1 次提交
- J
  add assign_value and hard_sigmoid, add fluid_type (#1983) · 92eeabeb
  由 juncaipeng 提交于 9月 09, 2019
```
* add assign_value op, arm kernel and test, add fluid_type, test=develop

* add hard_sigmoid, test=develop
```
  92eeabeb
07 9月, 2019 1 次提交

add lite x86 ops for ASR test=develop (#1981) · 7014a76b

由 lijianshe02 提交于 9月 07, 2019

* add lite x86 ops for ASR test=develop

* add lite x86 ops for ASR test=develop

* fix x86 ci run test problems test=develop

* fix mkl path for CI test=develop

7014a76b

06 9月, 2019 2 次提交

W
modify nearest_interpolate when attr align_corners=false (bug_fix) (#1969) · 4a3a45b6
由 Wilber 提交于 9月 06, 2019
```
* modify slice op and add slice test

* modify nearest_polate when align_corners=false (bugfix)
```
4a3a45b6

add cudnn conv fp32, int8 support (#1974) · f3124b30

由 Zhaolong Xing 提交于 9月 06, 2019

* paddle lite cuda init
can run model with leaky_relu

* add the missing file.
test=develop

* add the load from memory interface.
test=develop

* refine this pr. fix comments
fix ci error
test=develop

* conv impl
fp32:
conv, conv+bais, conv+bias+relu, conv+bias+leaky_relu

int8:
conv, conv+bais+relu(int8 or fp32 output), conv+bias+leaky_relu(int8 or fp32 output)

can run conv+ bias+relu using cxx_api
test=develop

* move the lite/cuda/math to backends/cuda/math
test=develop

f3124b30

04 9月, 2019 1 次提交
- W
  modify slice op and add slice test (#1944) · cfc7af76
  由 Wilber 提交于 9月 04, 2019
```
* modify slice op and add slice test

* modify slice op bug
```
  cfc7af76
03 9月, 2019 2 次提交
- H
  
  move npu into backends(directory) and move python/ into tools/python (#1958) · c5e65402
  由 huzhiqiang 提交于 9月 03, 2019
  
  c5e65402
- H
  
  create backends directory and move hardware backends into it (#1954) · 31ee212a
  由 huzhiqiang 提交于 9月 03, 2019
  
  31ee212a