提交 · ac1b2f9f10c19e1cc88ab122b01bb070e5222fb6 · PaddlePaddle / Paddle-Lite

28 10月, 2019 1 次提交

[LITE][XPU] initial support for XPU (#2202) · ac1b2f9f

由 hong19860320 提交于 10月 28, 2019

* Initial support for XPU
* Fix compiling errors of XPU
* Move XPU op kernel bridges from backends to kernels to fix deps order
* Change the namespace and directory of XPU bridges
* Add XPU SDK
* Fix header files and namespace of XPU SDK
* Add unit tests for relu and conv2d ops
* Restore the modification of paddle_api_test
* Supports simple model which contains only a relu layer
* Add compiling scripts for XPU
* Fix compiling errors of XPU
* Add comments for XPU LoadModel and BuildModel

ac1b2f9f

27 10月, 2019 1 次提交

model dynamic library tailoring (#2256) · b16917a4

由 huzhiqiang 提交于 10月 27, 2019

* add shell file to automatically build and collect publish result test=develop
* modify API inference of model_optimize_tool and add option for tiny&full publish test=develop

b16917a4

26 10月, 2019 1 次提交

Fix conv_bn fuser with no elemwise op added, Fix conv_elem with original conv... · e914f0da

由 Yuan Shuai 提交于 10月 26, 2019

Fix conv_bn fuser with no elemwise op added, Fix conv_elem with original conv with conv_bias (#2211)

* Fix conv_bn fuser with no elemwise op added. test=develop

* fix match not bug. test=develop

* Fix conv-bn pass. test=develop

* Fix conv-bn pass. test=develop

* Fix conv-bn pass. test=develop

* Fix conv-bn pass. test=develop

* Fix conv-bn pass. test=develop

* Fix conv-bn fuse pass. test=develop

* Fix conv-bn fuser consider the case: enable_int8=true without conv_bias. test=develop

* Fix back, consider enable_int8. test=develop

* Fix Bug for conv-bn quant pass. test=develop

* Fix conv-elemwise considering origin conv with conv_bias. test=develop

* Fix code format. test=develop

* simplify. test=develop

* simplify. test=develop

* Fix conv_elem. test=develop

e914f0da

25 10月, 2019 1 次提交
- H
  X86 dynamic compile (#2218) · 68003c0e
  由 huzhiqiang 提交于 10月 25, 2019
```
build x86 dynamic library in ci_build.sh
./lite/tool/ci_build.sh build_test_server
```
  68003c0e
24 10月, 2019 3 次提交

Make inceptionv4, resnet50, googlenet can run on x86 paltform (#2250) · ca7fefa1

由 liu zhengxi 提交于 10月 24, 2019

* make inceptionv4, resnet50, googlenet can run on x86 paltform and fix the compare part in x86 unittests, test=develop

* fix googlenet tests for benchmark record, test=develop

* [framework][profile] fix profile dump bug when op is feed and fetch test=develop (sangoly)

ca7fefa1

Fix gemv int8 error (#2249) · 3ea368b1

由 Xiaoyang LI 提交于 10月 24, 2019

* remove log in reshape, fix conv error when padding size=4, test=develop

* fix style, test=develop

* remove useless code, test=develop

* remove redundant model test file, test=develop

* change cluster to power_mode, test=develop

* fix build error, test=develop

* change cluster to power_mode, test=develop

* change opt_nb to use_optimize_nb, test=develop

* null, test=develop

* add gemv-int8 test, fix clang build error, test=develop

* fix gemv-int8 error when build with clang, test=develop

3ea368b1

S
[python api] add armlinux supported and publish paddle-lite python (#2252) · 60709e60
由 sangoly 提交于 10月 24, 2019
```
* [python api] add armlinux supported and publish paddle-lite python lib & demo

* add cuda build
test=develop
```
60709e60

23 10月, 2019 8 次提交

W
modify yolobox_cuda to support multiple runs (#2245) · e4b113eb
由 Wilber 提交于 10月 23, 2019
```
* modify yolobox_cuda to support multiple runs test=develop
```
e4b113eb
T
fix: fix fc batch bug (#2244) · ea4a5854
由 TianXiaogang 提交于 10月 23, 2019
```
* fix: fix fc batch bug

* fix:fix fc bug;test=develop
```
ea4a5854
石
Add kernel version table and update framework.proto, test=develop (#2243) · 842326f7
由石晓伟提交于 10月 23, 2019
```
* update framework.proto

* add compatibility check, test=develop

* remove head files, test=develop
```
842326f7

remove log in reshape, fix conv error when padding size=4 (#2199) · f8ff5aa4

由 Xiaoyang LI 提交于 10月 23, 2019

* remove log in reshape, fix conv error when padding size=4, test=develop

* fix style, test=develop

* remove useless code, test=develop

* remove redundant model test file, test=develop

* change cluster to power_mode, test=develop

* fix build error, test=develop

* change cluster to power_mode, test=develop

* change opt_nb to use_optimize_nb, test=develop

* null, test=develop

f8ff5aa4

Enable pool2d, dropout, transpose and transpose2 op on x86 (#2226) · 8006f5e6

由 liu zhengxi 提交于 10月 23, 2019

* enable pool2d op on x86 and add its unit tests, test=develop

* enable dropout op and add its unit tests, test=develop

* add tranpose, transpose2 op and add their unit tests, test=develop

8006f5e6

S
add python api (#2225) · bbfaacec
由 sangoly 提交于 10月 23, 2019
```
* [python api] init add python api test=develop
```
bbfaacec
X
fix python3 build optimize tool error · 05e96136
由 Xiaoyang LI 提交于 10月 23, 2019
```
* fix python3 build error, test=develop

* fix ci, test=develop
```
05e96136
G
Make armlinux compile libpaddle_light_api_shared.so (#2238) · 0ab6cf42
由 guofei 提交于 10月 23, 2019
```
* Make armlinux compile libpaddle_light_api_shared.so
test=develop
```
0ab6cf42

22 10月, 2019 4 次提交

Modify parse_op_registry (#2239) · 33601adf

由 juncaipeng 提交于 10月 22, 2019

* modify parse_op_registry. When REGISTER_LITE_OP and op_name not in the same row, it also can obtain op_name, test=develop

33601adf

Optimize quant_dequant (#2215) · aefb4ea3

由 juncaipeng 提交于 10月 22, 2019

* Add DeleteQuantOpFuser
* Add fake_quantize_dequantize_moving_avg_abs_max_op
* Add DeleteQuantDequantOpFuser

aefb4ea3

Z
remove feed and fetch for npu subgraph pass (#2230) · 52a093ee
由 zhupengyang 提交于 10月 22, 2019
```
test=develop
```
52a093ee

Transformer pr (#2214) · 330644b0

由 TianXiaogang 提交于 10月 22, 2019

* feat: add beam_search_special function for support nlp model

* fix: add beam_search_compute kernel input and output

* feat: add assign op & copy_compute kernel

* feat: add fill_const_batch_size_like op & kernel

* feat: add layer_norm op and kernel and ut

* fix: fix some bugs
    fix mul_op infer_shape bug when x_dim_idx = 2, x_dims.size()=3 & y_dim_idx = 1, y_dims.size()=2
    fix elementwise_compute bug when y axis is all 1
    fix beam_search choose math_func wrong bug
    fix layer_norm get attr bug
    fix fill_constant_batch_size_like shape_set bug

* feat: add gather op and kernel & and transform ut

* feats: add ops and fix bugs to support transformer op
       fix type_cast passes to skip `while`
       fix elementwise infer_shape bug when x.dims=3 and y.dims={1} & axis=0
       fix lookup_table compute bug
       fix read_from_array/beam_search/increment/compate/gather ops data_type problems

* fix:
    transfomer ut add word read inferface
    fix copy/gather/norm/layer_norm include path problem

* fix:debug info

* fix: fix input reshape bug

* fix: fix norm bug

* style: style fix & test=develop

* style: fix operators cmakelist

* style: fix operators cmakelist; test=develop

* fix and test=develop

* fix and test=develop

* style: style fix; test=develop

330644b0

21 10月, 2019 5 次提交
- 石
  link static library with cuda, test=develop (#2228) · 31f1e382
  由石晓伟提交于 10月 21, 2019
```
* add static libraries of cuda, test=develop

* update cuda make
```
  31f1e382
- to support yolov3 unet alexnet can run on tx2 (#2216) · d3d6ed4b
  由 myq406450149 提交于 10月 21, 2019
```
* add gpu kernel mul pool relu scale softmax dropout bilinear_interp and can run in tx2

* rm GREATER_EQUAL
```
  d3d6ed4b
- Y
  add cuda op(pool & softmax), support conv with padding_algorithm · 1a18d682
  由 yiicy 提交于 10月 21, 2019
```
* cuda add softmax and pool op

* * fix armlinux can find sys/system_properties.h
* conv add padding_algorithm
test=develop

* delete padding_algorithm in op param, test=develop

* fix bugs, test=develop
```
  1a18d682
- X
  
  fix ios build error, test=develop (#2222) · aefe9323
  由 Xiaoyang LI 提交于 10月 21, 2019
  
  aefe9323
- H
  Fix ‘Large memory usage of Naive model loading’ (#2175) · b0f94214
  由 huzhiqiang 提交于 10月 21, 2019
```
Fix ‘Large memory usage of Naive model loading’  (#2175)
```
  b0f94214
18 10月, 2019 2 次提交

W
fix yolobox_cuda_test (#2208) · 2f57f5b4
由 Wilber 提交于 10月 18, 2019
```
fix yolobox_cuda test precision error
```
2f57f5b4

Fix codestyle of GetInputName&GetOutputName (#2185) · 27a40b8f

由 huzhiqiang 提交于 10月 18, 2019

* add shell file to automatically build and collect publish result test=develop

* modify codestyle of getInputNames test=develop

* test=develop

* rm publish.sh

* remove copy of func param

* test=develop

* test=devcelop

* test=develop

* test=develop

* const & test=develop

* modify variable defination test=develop

* test=develop

* test=develop

* test=develop

* test=develop

27a40b8f

17 10月, 2019 4 次提交

J

add bilinear_interp_cuda_op, test=develop (#2197) · cb6b1b1c
由 juncaipeng 提交于 10月 17, 2019

cb6b1b1c

fix npu path (#2210) · c10a8b16

由 zhupengyang 提交于 10月 17, 2019

* move lite/backends/npu/bridges --> lite/kernels/npu/

test=develop

* fix namespace for npu

test=develop

* mv npu runtime file to lite/backends/npu

test=develop

c10a8b16

speedup fp32 depthwise conv · c5cd78ab

由 HappyAngel 提交于 10月 17, 2019

* update con_dw

* update

* add conv_depthwise_3x3s1.cc and conv_depthwise_3x3s2.cc

* add conv_depthwise_3x3s1_fp32 and conv_depthwise_3x3s2_fp32

* add new conv_dw

* only support conv_dw pad=0, 1

* add conv_dw_s1 conv_dw_s2 fp32

*     //conv2_func _impl2{nullptr};
update conv_dw, add conv_3x3s1 and conv_3x3s2, pad=[0,1]

* fix format, test=develop

* fix formmat, test=develop

c5cd78ab

L

enable batch_norm op and add its unit tests, test=develop (#2201) · 95372548
由 liu zhengxi 提交于 10月 17, 2019

95372548

16 10月, 2019 3 次提交
- Z
  Ban feed and fetch op during inference (#2198) · ad541652
  由 Zhaolong Xing 提交于 10月 16, 2019
```
* init: delete feed and fetch op, using zero copy
test=develop

* delete the unused test
test=develop
```
  ad541652
- L
  enable conv2d op and its unit tests, test=develop (#2200) · 0fed350e
  由 liu zhengxi 提交于 10月 16, 2019
```
enable conv2d op and its unit tests on x86 device
```
  0fed350e
- S
  [framework][place] remove prefered_place and kHost in valid_places (#2192) · 17833acb
  由 sangoly 提交于 10月 16, 2019
```
* [framework][place] remove prefered_place, use place order in valid_place array instead test=develop

* remove kHost from valid_places test=develop
```
  17833acb
15 10月, 2019 5 次提交

J
Fix quant dequant fuse pass (#2190) · 31ab471e
由 juncaipeng 提交于 10月 15, 2019
```
* fix bug for accessing the removed node, test=develop
```
31ab471e
J
fix benchmark, test=develop (#2188) · dbb660ee
由 juncaipeng 提交于 10月 15, 2019
```
* fix benchmark, test=develop
```
dbb660ee
石

fix pass selection, test=develop (#2187) · a1a22a69
由石晓伟提交于 10月 15, 2019

a1a22a69

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of... · 26d05370

由 Yuan Shuai 提交于 10月 15, 2019

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of CONVERT_TYPE_TO and READ/WRITE image, memory reuse in ResetLazyImage2D (#2170)

* add macro of CONVERT_TYPE_TO and READ/WRITE image. test=develop

* add data type control. test=develop

* fix io op as general layout and precision. test=develop

* Fix memory reuse strategy for opencl image2d. test=develop

* remove std::array, std::map in about opencl backend. test=develop

26d05370

[NPU] Fix and refine the supporting of multi NPU models (#2037) · e184d474

由 hong19860320 提交于 10月 15, 2019

* [NPU] Fix the bug of loading multi NPU models
test=develop

* [NPU] Use lite tensor to store NPU model, fix the management of multi NPU models, support loading NPU model from memory and reduce the modification of framework
test=develop

* [NPU] Remove redundant header files for NPU bridges,
test=develop

* [NPU] fix NPU deps
test=develop

* [NPU] refine the compiling script for NPU
test=develop

* [NPU] remove redundant subdirectory in lite/CMakeLists.txt
test=develop

* [NPU] Fix and refine NPU test case
test=develop

* [NPU] revoke the modification of other non-NPU modules
test=develop

* [NPU] Remove NPU bridges if target is tiny publish
test=develop

e184d474

14 10月, 2019 2 次提交
- J
  fix bug for reshape op, test=develop (#2141) · 4a061a27
  由 juncaipeng 提交于 10月 14, 2019
```
* fix bug for reshape op, test=develop
```
  4a061a27
- Z
  align yolov3 cuda int8 (#2183) · ed38d79b
  由 Zhaolong Xing 提交于 10月 14, 2019
```
test=develop
```
  ed38d79b