提交 · d87509668a6a5d2a6ca18450c05c48c43d052dcf · PaddlePaddle / Paddle-Lite

10 12月, 2019 2 次提交

W
fix type_target_cast pass. support only copy once for multiple use arg. test=develop (#2572) · 8903c795
由 Wilber 提交于 12月 10, 2019
```
For multiple-use parameters, only copy once
```
8903c795

modify static_kernel_pass to support select the kernel according to input type (#2488) · 7ef0e7fe

由 Wilber 提交于 12月 10, 2019

修改了选kernel的逻辑，默认从模型文件中读取出lod_tensor的data type，在static_kernel_pick pass中如果kernel输入输出的类型与读取的data type完全一致，则选择该Kernel的概率增大。

- 增加 从模型文件__model__读取lod_tensor的data type到cpp::vardesc

- program中增加unordered_map<string, type>字段，并在 Program::PrepareWorkspace中对该字段赋值

- 修改了node.h文件，将const Type* 更改为Type*，并在SSAGraph::Build过程中为符合条件的type*赋值

- static_kernel_pick_pass中添加新规则，如果kernel的输入类型输出类型与__model__中存储的类型的一致，则score*=2。

- 支持模型中用到sequence_reverse_float kernel（输入输出均为float）和sequence_reverse_int64 kernel（输入输出均为int64），能够根据输入输出type选kernel

7ef0e7fe

03 12月, 2019 1 次提交
- Z
  fix quant dequant fuse pass bug (#2552) · 137d7a6d
  由 Zhaolong Xing 提交于 12月 03, 2019
```
test=develop
```
  137d7a6d
29 11月, 2019 2 次提交
- Y
  [LITE][PASS] Fix static kernel pick pass, if op is not int8, but kernel is... · ddce609e
  由 Yuan Shuai 提交于 11月 29, 2019
```
[LITE][PASS] Fix static kernel pick pass, if op is not int8, but kernel is int8. test=develop (#2526)
```
  ddce609e
- Z
  [NPU] add reduce_mean op bridge and unit test (#2522) · c809321d
  由 zhupengyang 提交于 11月 29, 2019
```
* [NPU] add reduce_mean op bridge and unit test

test=develop

* refine xpu_pass place order; add bridges use

test=develop
```
  c809321d
26 11月, 2019 1 次提交

[XPU][NPU] update conv and conv_transpose padding (#2493) · eb9a0238

由 zhupengyang 提交于 11月 26, 2019

* [XPU] update conv padding

test=develop

* [XPU] exclude xpu for passes

test=develop

* fix conv padding compute

test=develop

* [NPU] update 2-pad to 4-pad for conv and conv_transpose

test=develop

* reuse UpdatePadding code

test=develop

eb9a0238

22 11月, 2019 1 次提交
- H
  [LITE][ALL] Refine NPU and XPU passes, fix the pass matching based on the... · c62fd634
  由 hong19860320 提交于 11月 22, 2019
```
[LITE][ALL] Refine NPU and XPU passes, fix the pass matching based on the bound targets and excluded targets (#2477)
```
  c62fd634
18 11月, 2019 1 次提交

[LITE][OPENCL] Enable full and light api for OpenCL (#2331) · d242bdfb

由 Yuan Shuai 提交于 11月 18, 2019

* Fix bug target for kHost and kARM not equal. test=develop

* Fix license. test=develop

* add debug -g option. test=develop

* enable opencl demo. test=develop

* Fix model_optimize_tool found no opencl kernel. test=develop

* add more vlog. test=develop

* remove macro LITE_WITH_OPENCL, LITE_WITH_FPGA in passes. test=develop

* Fix valid_places in mobilenetv1_test. test=develop

* Fix bug of find no real output of fetch, after tool OPs of optimzer passes. test=develop

* Fix vlog as log message in model_optimize_tool. test=develop

* fix miscs. test=develop

* fix comment. test=develop

* Fix misspell of opencl, fpga kernels name in lite/api/CMakeLists.txt. test=develop

* add opencl macro in full_api of demo. test=develop

d242bdfb

08 11月, 2019 1 次提交

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the... · aacba6f5

由 hong19860320 提交于 11月 08, 2019

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the limitation of the dimensions of input tensors in the graph compute kernel, and refine log message (#2395)

test=develop

aacba6f5

06 11月, 2019 1 次提交
- J
  add channel_wise_dequantized_max_abs op and ChannelWiseDequantOpFuser (#2368) · ba85799c
  由 juncaipeng 提交于 11月 06, 2019
```
* add channel_wise_dequantized_max_abs op and ChannelWiseDequantOpFuser, test=develop
```
  ba85799c
05 11月, 2019 2 次提交
- Z
  [XPU] not run x86 kernels for xpu (#2373) · 7e27b7bf
  由 zhupengyang 提交于 11月 05, 2019
```
test=develop
```
  7e27b7bf
- L
  fix StepRNN model run related bugs (#2300) · 7f5c0ca1
  由 lijianshe02 提交于 11月 05, 2019
```
* fix step rnn model run bugs test=develop
```
  7f5c0ca1
01 11月, 2019 1 次提交
- 石
  
  refactor: BindTargets and ExcludeTargets, test=develop (#2321) · 92179da1
  由石晓伟提交于 11月 01, 2019
  
  92179da1
31 10月, 2019 1 次提交
- Y
  [BugFix] Fix conv bn bug, check bias in pass (#2313) · c810ad2e
  由 Yuan Shuai 提交于 10月 31, 2019
```
* Fix conv bn bug. test=develop

* Fix bug in pattern_matcher. test=develop
```
  c810ad2e
30 10月, 2019 2 次提交
- H
  
  [LITE][NPU] Use FullConnection op to solve the compatibility between Kirin 810 and 990 (#2283) · f3e5d3e5
  由 hong19860320 提交于 10月 30, 2019
  
  f3e5d3e5
- X
  
  turn off conv_leaky_relu fusion when target is not cuda · 756140a8
  由 Xiaoyang LI 提交于 10月 30, 2019
  
  756140a8
29 10月, 2019 2 次提交
- Y
  Fix target bug: kHost and kARM not equal. test=develop (#2273) · 02888e52
  由 Yuan Shuai 提交于 10月 29, 2019
```
* Fix bug target for kHost and kARM not equal. test=develop

* Fix license. test=develop
```
  02888e52
- H
  [LITE][NPU] Add supporting for Huawei offical DDK (#2262) · dc2b853e
  由 hong19860320 提交于 10月 29, 2019
```
* Add supporting for Huawei offical DDK
* Fix the param of graph op in NPU graph computing kernel
```
  dc2b853e
28 10月, 2019 1 次提交

[LITE][XPU] initial support for XPU (#2202) · 06d058fe

由 hong19860320 提交于 10月 28, 2019

* Initial support for XPU
* Fix compiling errors of XPU
* Move XPU op kernel bridges from backends to kernels to fix deps order
* Change the namespace and directory of XPU bridges
* Add XPU SDK
* Fix header files and namespace of XPU SDK
* Add unit tests for relu and conv2d ops
* Restore the modification of paddle_api_test
* Supports simple model which contains only a relu layer
* Add compiling scripts for XPU
* Fix compiling errors of XPU
* Add comments for XPU LoadModel and BuildModel

06d058fe

26 10月, 2019 1 次提交

Fix conv_bn fuser with no elemwise op added, Fix conv_elem with original conv... · 74a6980e

由 Yuan Shuai 提交于 10月 26, 2019

Fix conv_bn fuser with no elemwise op added, Fix conv_elem with original conv with conv_bias (#2211)

* Fix conv_bn fuser with no elemwise op added. test=develop

* fix match not bug. test=develop

* Fix conv-bn pass. test=develop

* Fix conv-bn pass. test=develop

* Fix conv-bn pass. test=develop

* Fix conv-bn pass. test=develop

* Fix conv-bn pass. test=develop

* Fix conv-bn fuse pass. test=develop

* Fix conv-bn fuser consider the case: enable_int8=true without conv_bias. test=develop

* Fix back, consider enable_int8. test=develop

* Fix Bug for conv-bn quant pass. test=develop

* Fix conv-elemwise considering origin conv with conv_bias. test=develop

* Fix code format. test=develop

* simplify. test=develop

* simplify. test=develop

* Fix conv_elem. test=develop

74a6980e

22 10月, 2019 3 次提交

Optimize quant_dequant (#2215) · f480d474

由 juncaipeng 提交于 10月 22, 2019

* Add DeleteQuantOpFuser
* Add fake_quantize_dequantize_moving_avg_abs_max_op
* Add DeleteQuantDequantOpFuser

f480d474

Z
remove feed and fetch for npu subgraph pass (#2230) · 4e05ea29
由 zhupengyang 提交于 10月 22, 2019
```
test=develop
```
4e05ea29

Transformer pr (#2214) · f0a6c1eb

由 TianXiaogang 提交于 10月 22, 2019

* feat: add beam_search_special function for support nlp model

* fix: add beam_search_compute kernel input and output

* feat: add assign op & copy_compute kernel

* feat: add fill_const_batch_size_like op & kernel

* feat: add layer_norm op and kernel and ut

* fix: fix some bugs
    fix mul_op infer_shape bug when x_dim_idx = 2, x_dims.size()=3 & y_dim_idx = 1, y_dims.size()=2
    fix elementwise_compute bug when y axis is all 1
    fix beam_search choose math_func wrong bug
    fix layer_norm get attr bug
    fix fill_constant_batch_size_like shape_set bug

* feat: add gather op and kernel & and transform ut

* feats: add ops and fix bugs to support transformer op
       fix type_cast passes to skip `while`
       fix elementwise infer_shape bug when x.dims=3 and y.dims={1} & axis=0
       fix lookup_table compute bug
       fix read_from_array/beam_search/increment/compate/gather ops data_type problems

* fix:
    transfomer ut add word read inferface
    fix copy/gather/norm/layer_norm include path problem

* fix:debug info

* fix: fix input reshape bug

* fix: fix norm bug

* style: style fix & test=develop

* style: fix operators cmakelist

* style: fix operators cmakelist; test=develop

* fix and test=develop

* fix and test=develop

* style: style fix; test=develop

f0a6c1eb

21 10月, 2019 1 次提交
- to support yolov3 unet alexnet can run on tx2 (#2216) · 57d8e42e
  由 myq406450149 提交于 10月 21, 2019
```
* add gpu kernel mul pool relu scale softmax dropout bilinear_interp and can run in tx2

* rm GREATER_EQUAL
```
  57d8e42e
17 10月, 2019 1 次提交

fix npu path (#2210) · 7c722a37

由 zhupengyang 提交于 10月 17, 2019

* move lite/backends/npu/bridges --> lite/kernels/npu/

test=develop

* fix namespace for npu

test=develop

* mv npu runtime file to lite/backends/npu

test=develop

7c722a37

16 10月, 2019 1 次提交

[framework][place] remove prefered_place and kHost in valid_places (#2192) · 3012088b

由 sangoly 提交于 10月 16, 2019

* [framework][place] remove prefered_place, use place order in valid_place array instead test=develop

* remove kHost from valid_places test=develop

3012088b

15 10月, 2019 4 次提交

J
Fix quant dequant fuse pass (#2190) · 9cc7dfa8
由 juncaipeng 提交于 10月 15, 2019
```
* fix bug for accessing the removed node, test=develop
```
9cc7dfa8
石

fix pass selection, test=develop (#2187) · da55f674
由石晓伟提交于 10月 15, 2019

da55f674

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of... · 72c11758

由 Yuan Shuai 提交于 10月 15, 2019

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of CONVERT_TYPE_TO and READ/WRITE image, memory reuse in ResetLazyImage2D (#2170)

* add macro of CONVERT_TYPE_TO and READ/WRITE image. test=develop

* add data type control. test=develop

* fix io op as general layout and precision. test=develop

* Fix memory reuse strategy for opencl image2d. test=develop

* remove std::array, std::map in about opencl backend. test=develop

72c11758

[NPU] Fix and refine the supporting of multi NPU models (#2037) · 7a731b7f

由 hong19860320 提交于 10月 15, 2019

* [NPU] Fix the bug of loading multi NPU models
test=develop

* [NPU] Use lite tensor to store NPU model, fix the management of multi NPU models, support loading NPU model from memory and reduce the modification of framework
test=develop

* [NPU] Remove redundant header files for NPU bridges,
test=develop

* [NPU] fix NPU deps
test=develop

* [NPU] refine the compiling script for NPU
test=develop

* [NPU] remove redundant subdirectory in lite/CMakeLists.txt
test=develop

* [NPU] Fix and refine NPU test case
test=develop

* [NPU] revoke the modification of other non-NPU modules
test=develop

* [NPU] Remove NPU bridges if target is tiny publish
test=develop

7a731b7f

14 10月, 2019 2 次提交
- Z
  align yolov3 cuda int8 (#2183) · 80d35725
  由 Zhaolong Xing 提交于 10月 14, 2019
```
test=develop
```
  80d35725
- J
  Optimize quant_dequant_fuse_pass (#2169) · 253acb80
  由 juncaipeng 提交于 10月 14, 2019
```
* optimize quant_dequant_fuse_pass, test=develop
```
  253acb80
11 10月, 2019 2 次提交

CUDA: can run yolov3 int8 (#2172) · 7931104f

由 Zhaolong Xing 提交于 10月 11, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

* can run yolov3 int8 test=develop

7931104f

[LITE][OPENCL] support image2d type (#2158) · 77cdbdce

由 Yuan Shuai 提交于 10月 11, 2019

* [LITE][OPENCL] support image2d. test=develop

* add context changed with consider image*. test=develop

* add layout, relu image kernels. test=develop

* replace image_data with data, mutable_image_data with mutable_data, test=develop

* comment unused var. test=develop

* remove unused var. test=develop

77cdbdce

27 9月, 2019 1 次提交

can run yolov3 fp32 on cuda devices (#2092) · 3d6d744f

由 Zhaolong Xing 提交于 9月 27, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

3d6d744f

26 9月, 2019 1 次提交
- S
  
  [Fc Fusion] fix fc fusion duplicative arguments bug test=develop (#2135) · 1e6bb8d5
  由 sangoly 提交于 9月 26, 2019
  
  1e6bb8d5
23 9月, 2019 1 次提交
- W
  
  model_test add host place (#2109) · 8bee4e29
  由 Wilber 提交于 9月 23, 2019
  
  8bee4e29
21 9月, 2019 1 次提交
- W
  
  add yolo_box to mem_optimize_pass's unuse op set · 8b6976b2
  由 Wilber 提交于 9月 21, 2019
  
  8b6976b2
20 9月, 2019 2 次提交
- H
  fix compiling and fix code style (#2088) · c10bc6d6
  由 huzhiqiang 提交于 9月 20, 2019
```
* fix compiling and fix code style test=develop
```
  c10bc6d6
- Z
  1. the split op's bug will triger memory optimize pass failed. (#2070) · 977a66fc
  由 Zhaolong Xing 提交于 9月 20, 2019
```
test=develop
```
  977a66fc