提交 · 2e40b2f7843e1a7cdcd78eec12beaabc91ba9915 · PaddlePaddle / Paddle-Lite

23 10月, 2019 2 次提交
- J
  
  add cuda_kernels just with LITE_WITH_CUDA · 5cc5da6f
  由 juncaipeng 提交于 10月 23, 2019
  
  5cc5da6f
- G
  Make armlinux compile libpaddle_light_api_shared.so (#2238) · caeb9c82
  由 guofei 提交于 10月 23, 2019
```
* Make armlinux compile libpaddle_light_api_shared.so
test=develop
```
  caeb9c82
22 10月, 2019 3 次提交

J

support build cuda lib · 0da40cd6
由 juncaipeng 提交于 10月 22, 2019

0da40cd6

Optimize quant_dequant (#2215) · f480d474

由 juncaipeng 提交于 10月 22, 2019

* Add DeleteQuantOpFuser
* Add fake_quantize_dequantize_moving_avg_abs_max_op
* Add DeleteQuantDequantOpFuser

f480d474

Transformer pr (#2214) · f0a6c1eb

由 TianXiaogang 提交于 10月 22, 2019

* feat: add beam_search_special function for support nlp model

* fix: add beam_search_compute kernel input and output

* feat: add assign op & copy_compute kernel

* feat: add fill_const_batch_size_like op & kernel

* feat: add layer_norm op and kernel and ut

* fix: fix some bugs
    fix mul_op infer_shape bug when x_dim_idx = 2, x_dims.size()=3 & y_dim_idx = 1, y_dims.size()=2
    fix elementwise_compute bug when y axis is all 1
    fix beam_search choose math_func wrong bug
    fix layer_norm get attr bug
    fix fill_constant_batch_size_like shape_set bug

* feat: add gather op and kernel & and transform ut

* feats: add ops and fix bugs to support transformer op
       fix type_cast passes to skip `while`
       fix elementwise infer_shape bug when x.dims=3 and y.dims={1} & axis=0
       fix lookup_table compute bug
       fix read_from_array/beam_search/increment/compate/gather ops data_type problems

* fix:
    transfomer ut add word read inferface
    fix copy/gather/norm/layer_norm include path problem

* fix:debug info

* fix: fix input reshape bug

* fix: fix norm bug

* style: style fix & test=develop

* style: fix operators cmakelist

* style: fix operators cmakelist; test=develop

* fix and test=develop

* fix and test=develop

* style: style fix; test=develop

f0a6c1eb

21 10月, 2019 1 次提交
- 石
  link static library with cuda, test=develop (#2228) · 7c69b6b4
  由石晓伟提交于 10月 21, 2019
```
* add static libraries of cuda, test=develop

* update cuda make
```
  7c69b6b4
18 10月, 2019 1 次提交

Fix codestyle of GetInputName&GetOutputName (#2185) · 8591aaec

由 huzhiqiang 提交于 10月 18, 2019

* add shell file to automatically build and collect publish result test=develop

* modify codestyle of getInputNames test=develop

* test=develop

* rm publish.sh

* remove copy of func param

* test=develop

* test=devcelop

* test=develop

* test=develop

* const & test=develop

* modify variable defination test=develop

* test=develop

* test=develop

* test=develop

* test=develop

8591aaec

16 10月, 2019 2 次提交

Z
Ban feed and fetch op during inference (#2198) · 75e8a6fc
由 Zhaolong Xing 提交于 10月 16, 2019
```
* init: delete feed and fetch op, using zero copy
test=develop

* delete the unused test
test=develop
```
75e8a6fc

[framework][place] remove prefered_place and kHost in valid_places (#2192) · 3012088b

由 sangoly 提交于 10月 16, 2019

* [framework][place] remove prefered_place, use place order in valid_place array instead test=develop

* remove kHost from valid_places test=develop

3012088b

15 10月, 2019 2 次提交

J
fix benchmark, test=develop (#2188) · 4d530acc
由 juncaipeng 提交于 10月 15, 2019
```
* fix benchmark, test=develop
```
4d530acc

[NPU] Fix and refine the supporting of multi NPU models (#2037) · 7a731b7f

由 hong19860320 提交于 10月 15, 2019

* [NPU] Fix the bug of loading multi NPU models
test=develop

* [NPU] Use lite tensor to store NPU model, fix the management of multi NPU models, support loading NPU model from memory and reduce the modification of framework
test=develop

* [NPU] Remove redundant header files for NPU bridges,
test=develop

* [NPU] fix NPU deps
test=develop

* [NPU] refine the compiling script for NPU
test=develop

* [NPU] remove redundant subdirectory in lite/CMakeLists.txt
test=develop

* [NPU] Fix and refine NPU test case
test=develop

* [NPU] revoke the modification of other non-NPU modules
test=develop

* [NPU] Remove NPU bridges if target is tiny publish
test=develop

7a731b7f

14 10月, 2019 2 次提交
- H
  add GetInputNames 、 GetOutPutNames 、 GetInputByName and GetTensor method (#2154) · 56151776
  由 huzhiqiang 提交于 10月 14, 2019
```
* add GetInputNames and GetOutPutNames and GetInputByName method test=develop
```
  56151776
- J
  Optimize quant_dequant_fuse_pass (#2169) · 253acb80
  由 juncaipeng 提交于 10月 14, 2019
```
* optimize quant_dequant_fuse_pass, test=develop
```
  253acb80
11 10月, 2019 2 次提交
- J
  
  add rsqrt op, test=develop (#2176) · dfce4621
  由 juncaipeng 提交于 10月 11, 2019
  
  dfce4621
- H
  move the method of SetThread and SetPowerMode from MobileConfig into ConfigBase (#2147) · 1ae9239e
  由 huzhiqiang 提交于 10月 11, 2019
```
* move the method of SetThread and SetPowerMode from MobileConfig into ConfigBase 
* cxxPredictor will support SetThread and SetPowerMode method
```
  1ae9239e
27 9月, 2019 1 次提交

can run yolov3 fp32 on cuda devices (#2092) · 3d6d744f

由 Zhaolong Xing 提交于 9月 27, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

3d6d744f

25 9月, 2019 2 次提交
- X
  
  fix MobileConfig get mode and threads error (#2134) · 6b1c6faa
  由 Xiaoyang LI 提交于 9月 25, 2019
  
  6b1c6faa
- X
  
  fix arm device_info error, fix bind big core error, improve 855 performance, test=develop (#2133) · 10fbdc5e
  由 Xiaoyang LI 提交于 9月 25, 2019
  
  10fbdc5e
23 9月, 2019 1 次提交
- W
  
  model_test add host place (#2109) · 8bee4e29
  由 Wilber 提交于 9月 23, 2019
  
  8bee4e29
19 9月, 2019 4 次提交

石

add full_api_static target and fix building errors, test=develop (#2064) · eef7ea0f

由石晓伟提交于 9月 19, 2019

* add full_api_static target and fix building errors, test=develop

* fix build errors, test=develop

* fix code style, test=develop

* fix lite/model_parser/pb/var_desc.cc, test=develop

* fix building errors, test=develop

* modify lite/tools/debug/CMakeLists.txt, test=develop

eef7ea0f

Bug fix for model save and load (#1992) · 8efbdc66

由 TianXiaogang 提交于 9月 19, 2019

* fix: fix model parser and save bug

* style: delete debug code

* fix: fix light_predictor program run model with subblock bug

8efbdc66

modify dynamic library: libpaddle_cxx_api.so (#2057) · fcdb7081

由 huzhiqiang 提交于 9月 19, 2019

(1)modify tiny publish so to make it excutable
(2)modify the bug of compiling in armlinux
(3)change 3 full publish so into 2

fcdb7081

S

[Java API] add getVersion() interface to get c++ lib's version information test=develop (#2063) · 71b05c94
由 sangoly 提交于 9月 19, 2019

71b05c94

18 9月, 2019 1 次提交

fix bias quantize error && fix clang build error (#2049) · 81dffbe8

由 Xiaoyang LI 提交于 9月 18, 2019

* fix gemm_int8, gemv-int8 and conv-int8 math function, add float bias

* change conv impl

* neon int8 kernel support float bias

* arm compute kernel support float bias

* add math_test target

* add tensor utils for testing, fix sgemm ut error

* add gemm_int8 unit test, support float bias

* fix build script

* add conv compute unit test for arm

* fix build script, test=develop

* fix fp32 dw conv3x3s1, test=develop

* add fp32 dw conv3x3s1, test=develop

* add armv7 fp32 dw conv3x3s1, test=develop

* add fp32 depthwise conv3x3s2, test=develop

* fix fp32 conv3x3 depthwise build error, test=develop

* fix gemm_like conv trans weights error, test=develop

* fix int8 depthwise conv3x3 error, test=develop

* turn on all test for arm fp32 conv, test=develop

* fix int8 conv1x1 error

* fix int8 direct conv3x3s1 error, test=develop

* fix int8 direct conv3x3s2, test=develop

* turn on all test for arm int8 conv, test=develop

* fix int8 fc error, change mobilenetv1-int8 ground-truth result to fluid, test=develop

* remove debug info, strip ut binary, test=develop

* fix conv compute error, test=develop

* change Init() to ReInitWhenNeeded(), test=develop

* fix code style, test=develop

* remote engine_test, test=develop

* fix building server tests error, test=develop

* fix sdot clang build error, test=develop

* fix sgemm ut timeout error, test=develop

* fix clang build error, test=develop

* turn off math basic test due to ci time out, test=develop

* fix conv_int8 ut error, test=develop

81dffbe8

17 9月, 2019 1 次提交
- S
  [Cxx API] add build-in version info (#2047) · e7eea682
  由 sangoly 提交于 9月 17, 2019
```
* [Cxx API] add build-in version info

* update: add version.h.in template
```
  e7eea682
16 9月, 2019 1 次提交
- H
  
  add dynamic library for tiny and full publish (#2036) · 7e6a9030
  由 huzhiqiang 提交于 9月 16, 2019
  
  7e6a9030
13 9月, 2019 1 次提交
- Z
  
  Add the memory optim pass (#2018) · 89b3026c
  由 Zhaolong Xing 提交于 9月 13, 2019
  
  89b3026c
12 9月, 2019 2 次提交
- Y
  
  integrate model_optimize_tool compilation to build.sh (#2033) · 3d9028de
  由 Yan Chunwei 提交于 9月 12, 2019
  
  3d9028de
- W
  add unsqueeze and range op (x2paddle) (#1988) · 3c08f676
  由 Wilber 提交于 9月 12, 2019
```
* add unsqueeze and range op. modify concat op test=develop

* modify exception in range_test_x86
```
  3c08f676
11 9月, 2019 2 次提交
- Y
  
  make model_optimize_tool run on host (#1990) · 83d4b0e8
  由 Yan Chunwei 提交于 9月 11, 2019
  
  83d4b0e8
- L
  add slice op, reshape op, reshape2 op, squeeze op, squeeze2 op for x86 (#2005) · 13bbd2b8
  由 liu zhengxi 提交于 9月 11, 2019
```
add slice op, reshape op,  reshape2 op, squeeze op and squeeze2 op and their unittests for x86
```
  13bbd2b8
10 9月, 2019 3 次提交
- W
  
  add elementwise_sub and modify argmax (#1964) · 62ea82d0
  由 Wilber 提交于 9月 10, 2019
  
  62ea82d0
- J
  Modify detection test (#2000) · 111db475
  由 juncaipeng 提交于 9月 10, 2019
```
* add assign_value op, arm kernel and test, add fluid_type, test=develop

* add hard_sigmoid, test=develop

* use image and new imple to test detection model, delete faster_rcnn_test, test=develop
```
  111db475
- X
  fix model_optimize_tool error when using host kernel, fix reshape op build... · f25a4571
  由 Xiaoyang LI 提交于 9月 10, 2019
```
fix model_optimize_tool error when using host kernel, fix reshape op build error on ios, test=develop (#1984)
```
  f25a4571
09 9月, 2019 2 次提交

J
add assign_value and hard_sigmoid, add fluid_type (#1983) · 92eeabeb
由 juncaipeng 提交于 9月 09, 2019
```
* add assign_value op, arm kernel and test, add fluid_type, test=develop

* add hard_sigmoid, test=develop
```
92eeabeb

Add concat and elementwise_add cuda kernel (#1979) · 6d1da405

由 Pei Yang 提交于 9月 09, 2019

* add nearest_interp_cuda kernel, test=develop

* add concat op and elementwise_add op

* remove eigen dependency from nearest_interp cuda kernel, test=develop

* free cuda pointers, test=develop

6d1da405

06 9月, 2019 2 次提交

Z
add interpolate fuse pass (#1980) · c49958a2
由 zhupengyang 提交于 9月 06, 2019
```
test=develop
```
c49958a2

add cudnn conv fp32, int8 support (#1974) · f3124b30

由 Zhaolong Xing 提交于 9月 06, 2019

* paddle lite cuda init
can run model with leaky_relu

* add the missing file.
test=develop

* add the load from memory interface.
test=develop

* refine this pr. fix comments
fix ci error
test=develop

* conv impl
fp32:
conv, conv+bais, conv+bias+relu, conv+bias+leaky_relu

int8:
conv, conv+bais+relu(int8 or fp32 output), conv+bias+leaky_relu(int8 or fp32 output)

can run conv+ bias+relu using cxx_api
test=develop

* move the lite/cuda/math to backends/cuda/math
test=develop

f3124b30

03 9月, 2019 2 次提交

H

move npu into backends(directory) and move python/ into tools/python (#1958) · c5e65402
由 huzhiqiang 提交于 9月 03, 2019

c5e65402

rewrite multiclass_nms according to fluid, test=develop (#1945) · deaddf9d

由 juncaipeng 提交于 9月 03, 2019

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

* add box_clip op and fix tensor slice bug

* remove add four op twice

* rewrite the implement for box_coder and sequence_expand, add faster_rcnn_test, test=develop

* fix test bug of box_clip in x86 server, test=develop

* rewrite multiclass_nms according to fluid, test=develop

* fix param load bug in box_coder and multiclass_nms op, test=develop

* fix value transfor error in multiclass_nms, test=develop

deaddf9d