提交 · bbfaacecfc49d09fdbdbd9f04ef538324f96447b · PaddlePaddle / Paddle-Lite

22 10月, 2019 3 次提交

Modify parse_op_registry (#2239) · 33601adf

由 juncaipeng 提交于 10月 22, 2019

* modify parse_op_registry. When REGISTER_LITE_OP and op_name not in the same row, it also can obtain op_name, test=develop

33601adf

Optimize quant_dequant (#2215) · aefb4ea3

由 juncaipeng 提交于 10月 22, 2019

* Add DeleteQuantOpFuser
* Add fake_quantize_dequantize_moving_avg_abs_max_op
* Add DeleteQuantDequantOpFuser

aefb4ea3

Transformer pr (#2214) · 330644b0

由 TianXiaogang 提交于 10月 22, 2019

* feat: add beam_search_special function for support nlp model

* fix: add beam_search_compute kernel input and output

* feat: add assign op & copy_compute kernel

* feat: add fill_const_batch_size_like op & kernel

* feat: add layer_norm op and kernel and ut

* fix: fix some bugs
    fix mul_op infer_shape bug when x_dim_idx = 2, x_dims.size()=3 & y_dim_idx = 1, y_dims.size()=2
    fix elementwise_compute bug when y axis is all 1
    fix beam_search choose math_func wrong bug
    fix layer_norm get attr bug
    fix fill_constant_batch_size_like shape_set bug

* feat: add gather op and kernel & and transform ut

* feats: add ops and fix bugs to support transformer op
       fix type_cast passes to skip `while`
       fix elementwise infer_shape bug when x.dims=3 and y.dims={1} & axis=0
       fix lookup_table compute bug
       fix read_from_array/beam_search/increment/compate/gather ops data_type problems

* fix:
    transfomer ut add word read inferface
    fix copy/gather/norm/layer_norm include path problem

* fix:debug info

* fix: fix input reshape bug

* fix: fix norm bug

* style: style fix & test=develop

* style: fix operators cmakelist

* style: fix operators cmakelist; test=develop

* fix and test=develop

* fix and test=develop

* style: style fix; test=develop

330644b0

21 10月, 2019 1 次提交

add cuda op(pool & softmax), support conv with padding_algorithm · 1a18d682

由 yiicy 提交于 10月 21, 2019

* cuda add softmax and pool op

* * fix armlinux can find sys/system_properties.h
* conv add padding_algorithm
test=develop

* delete padding_algorithm in op param, test=develop

* fix bugs, test=develop

1a18d682

17 10月, 2019 1 次提交
- J
  
  add bilinear_interp_cuda_op, test=develop (#2197) · cb6b1b1c
  由 juncaipeng 提交于 10月 17, 2019
  
  cb6b1b1c
15 10月, 2019 1 次提交

[NPU] Fix and refine the supporting of multi NPU models (#2037) · e184d474

由 hong19860320 提交于 10月 15, 2019

* [NPU] Fix the bug of loading multi NPU models
test=develop

* [NPU] Use lite tensor to store NPU model, fix the management of multi NPU models, support loading NPU model from memory and reduce the modification of framework
test=develop

* [NPU] Remove redundant header files for NPU bridges,
test=develop

* [NPU] fix NPU deps
test=develop

* [NPU] refine the compiling script for NPU
test=develop

* [NPU] remove redundant subdirectory in lite/CMakeLists.txt
test=develop

* [NPU] Fix and refine NPU test case
test=develop

* [NPU] revoke the modification of other non-NPU modules
test=develop

* [NPU] Remove NPU bridges if target is tiny publish
test=develop

e184d474

14 10月, 2019 3 次提交
- J
  fix bug for reshape op, test=develop (#2141) · 4a061a27
  由 juncaipeng 提交于 10月 14, 2019
```
* fix bug for reshape op, test=develop
```
  4a061a27
- L
  fix asr modle related kernel bugs test=develop (#2179) · 2f035fec
  由 lijianshe02 提交于 10月 14, 2019
```
* fix asr modle related kernel bugs test=develop
```
  2f035fec
- J
  Optimize quant_dequant_fuse_pass (#2169) · 0260d322
  由 juncaipeng 提交于 10月 14, 2019
```
* optimize quant_dequant_fuse_pass, test=develop
```
  0260d322
11 10月, 2019 2 次提交

J

add rsqrt op, test=develop (#2176) · 78ddd64d
由 juncaipeng 提交于 10月 11, 2019

78ddd64d

CUDA: can run yolov3 int8 (#2172) · 29f448c6

由 Zhaolong Xing 提交于 10月 11, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

* can run yolov3 int8 test=develop

29f448c6

27 9月, 2019 1 次提交

can run yolov3 fp32 on cuda devices (#2092) · c4b5e32c

由 Zhaolong Xing 提交于 9月 27, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

c4b5e32c

20 9月, 2019 1 次提交
- Z
  1. the split op's bug will triger memory optimize pass failed. (#2070) · b7f5d94b
  由 Zhaolong Xing 提交于 9月 20, 2019
```
test=develop
```
  b7f5d94b
19 9月, 2019 2 次提交

石

add full_api_static target and fix building errors, test=develop (#2064) · 4a948cfc

由石晓伟提交于 9月 19, 2019

* add full_api_static target and fix building errors, test=develop

* fix build errors, test=develop

* fix code style, test=develop

* fix lite/model_parser/pb/var_desc.cc, test=develop

* fix building errors, test=develop

* modify lite/tools/debug/CMakeLists.txt, test=develop

4a948cfc

fix building model_optimize_tool error on mac (#2075) · 26925ab9

由 Xiaoyang LI 提交于 9月 19, 2019

* fix building model_optimize_tool error on mac, test=develop

* fix model_optimize_tool build error, test=develop

26925ab9

18 9月, 2019 1 次提交

fix bias quantize error && fix clang build error (#2049) · 8d6f475e

由 Xiaoyang LI 提交于 9月 18, 2019

* fix gemm_int8, gemv-int8 and conv-int8 math function, add float bias

* change conv impl

* neon int8 kernel support float bias

* arm compute kernel support float bias

* add math_test target

* add tensor utils for testing, fix sgemm ut error

* add gemm_int8 unit test, support float bias

* fix build script

* add conv compute unit test for arm

* fix build script, test=develop

* fix fp32 dw conv3x3s1, test=develop

* add fp32 dw conv3x3s1, test=develop

* add armv7 fp32 dw conv3x3s1, test=develop

* add fp32 depthwise conv3x3s2, test=develop

* fix fp32 conv3x3 depthwise build error, test=develop

* fix gemm_like conv trans weights error, test=develop

* fix int8 depthwise conv3x3 error, test=develop

* turn on all test for arm fp32 conv, test=develop

* fix int8 conv1x1 error

* fix int8 direct conv3x3s1 error, test=develop

* fix int8 direct conv3x3s2, test=develop

* turn on all test for arm int8 conv, test=develop

* fix int8 fc error, change mobilenetv1-int8 ground-truth result to fluid, test=develop

* remove debug info, strip ut binary, test=develop

* fix conv compute error, test=develop

* change Init() to ReInitWhenNeeded(), test=develop

* fix code style, test=develop

* remote engine_test, test=develop

* fix building server tests error, test=develop

* fix sdot clang build error, test=develop

* fix sgemm ut timeout error, test=develop

* fix clang build error, test=develop

* turn off math basic test due to ci time out, test=develop

* fix conv_int8 ut error, test=develop

8d6f475e

17 9月, 2019 3 次提交
- W
  modify norm kernel to run caffe_facedetection model (#2008) · 133a40d2
  由 Wilber 提交于 9月 17, 2019
```
* modify norm kernel to run caffe_facedetection model

* reserve bind norm, remove calc of norm output
```
  133a40d2
- L
  add fill_constant_batch_size_like op and add its unittest (#2044) · b91850b9
  由 liu zhengxi 提交于 9月 17, 2019
```
* add fill_constant_batch_size_like op and add its unittest
```
  b91850b9
- J
  fix yolo_box bug (#2034) · 7497a7d2
  由 juncaipeng 提交于 9月 17, 2019
```
* fix yolo_box bug, test=develop

* fix test bug for yolo_box, test=develop
```
  7497a7d2
16 9月, 2019 1 次提交
- L
  Gru op (#2002) · 1cb36af6
  由 lhl960107 提交于 9月 16, 2019
```
* add x86 gru&&relu&&sequence_expand_as op test=develop
```
  1cb36af6
12 9月, 2019 2 次提交
- W
  add unsqueeze and range op (x2paddle) (#1988) · cca0aec6
  由 Wilber 提交于 9月 12, 2019
```
* add unsqueeze and range op. modify concat op test=develop

* modify exception in range_test_x86
```
  cca0aec6
- W
  
  add min_max_aspect_ratios_order attr in prior box op test=develop (#2016) · f04ed39c
  由 Wilber 提交于 9月 12, 2019
  
  f04ed39c
11 9月, 2019 1 次提交
- L
  add slice op, reshape op, reshape2 op, squeeze op, squeeze2 op for x86 (#2005) · 909250b4
  由 liu zhengxi 提交于 9月 11, 2019
```
add slice op, reshape op,  reshape2 op, squeeze op and squeeze2 op and their unittests for x86
```
  909250b4
10 9月, 2019 2 次提交
- W
  
  add elementwise_sub and modify argmax (#1964) · 192320c4
  由 Wilber 提交于 9月 10, 2019
  
  192320c4
- X
  fix model_optimize_tool error when using host kernel, fix reshape op build... · f4826ac9
  由 Xiaoyang LI 提交于 9月 10, 2019
```
fix model_optimize_tool error when using host kernel, fix reshape op build error on ios, test=develop (#1984)
```
  f4826ac9
09 9月, 2019 1 次提交
- J
  add assign_value and hard_sigmoid, add fluid_type (#1983) · 9796c57d
  由 juncaipeng 提交于 9月 09, 2019
```
* add assign_value op, arm kernel and test, add fluid_type, test=develop

* add hard_sigmoid, test=develop
```
  9796c57d
07 9月, 2019 1 次提交

add lite x86 ops for ASR test=develop (#1981) · 25b775d6

由 lijianshe02 提交于 9月 07, 2019

* add lite x86 ops for ASR test=develop

* add lite x86 ops for ASR test=develop

* fix x86 ci run test problems test=develop

* fix mkl path for CI test=develop

25b775d6

06 9月, 2019 2 次提交

H
modify reshape2 OP test=dvelop (#1963) · 49052cb7
由 huzhiqiang 提交于 9月 06, 2019
```
modify reshape2 OP to add shape_tensor input
```
49052cb7

add cudnn conv fp32, int8 support (#1974) · 23d83c04

由 Zhaolong Xing 提交于 9月 06, 2019

* paddle lite cuda init
can run model with leaky_relu

* add the missing file.
test=develop

* add the load from memory interface.
test=develop

* refine this pr. fix comments
fix ci error
test=develop

* conv impl
fp32:
conv, conv+bais, conv+bias+relu, conv+bias+leaky_relu

int8:
conv, conv+bais+relu(int8 or fp32 output), conv+bias+leaky_relu(int8 or fp32 output)

can run conv+ bias+relu using cxx_api
test=develop

* move the lite/cuda/math to backends/cuda/math
test=develop

23d83c04

03 9月, 2019 2 次提交

Z
enhance interpolate op when there is no "scale" (#1957) · 64b4cf28
由 zhupengyang 提交于 9月 03, 2019
```
test=develop
```
64b4cf28

rewrite multiclass_nms according to fluid, test=develop (#1945) · 664e19cc

由 juncaipeng 提交于 9月 03, 2019

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

* add box_clip op and fix tensor slice bug

* remove add four op twice

* rewrite the implement for box_coder and sequence_expand, add faster_rcnn_test, test=develop

* fix test bug of box_clip in x86 server, test=develop

* rewrite multiclass_nms according to fluid, test=develop

* fix param load bug in box_coder and multiclass_nms op, test=develop

* fix value transfor error in multiclass_nms, test=develop

664e19cc

02 9月, 2019 1 次提交

Add ops and fix bugs for Faster RCNN (#1942) · cfd5abe5

由 juncaipeng 提交于 9月 02, 2019

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

* add box_clip op and fix tensor slice bug

* remove add four op twice

* rewrite the implement for box_coder and sequence_expand, add faster_rcnn_test, test=develop

* fix test bug of box_clip in x86 server, test=develop

cfd5abe5

01 9月, 2019 1 次提交

[ARM][CPU] Fix time counter of arm cpu profiler (#1925) · c096de0e

由 Yuan Shuai 提交于 9月 01, 2019

* Fix timer of arm cpu profiler. test=develop

* Fix un-added op in cmake.test=develop

* fix cmake error

* fix cmake error, test=develop

* Fix pass sequence. test=develop

* replace option with lite_option. test=develop

* disable profile mode by default. test=develop

* Fix error option name. test=develop

c096de0e

30 8月, 2019 1 次提交

Modify the slice op, stack op, reduce_mean op (#1921) · 440aae7a

由 liu zhengxi 提交于 8月 30, 2019

* add stack op and add reduce_mean op and their unit tests

* modify stack op output name and modify the for loop in reduce_mean op

* add HasAttr for slice op

440aae7a

29 8月, 2019 2 次提交

L

add stack op and add reduce_mean op and their unit tests (#1888) · 8ccd01a6
由 liu zhengxi 提交于 8月 29, 2019

8ccd01a6

ad ops for faster rcnn, including affine_channel, anchor_generator,... · f3035827

由 juncaipeng 提交于 8月 29, 2019

ad ops for faster rcnn, including affine_channel, anchor_generator, generate_proposals and roi_align (#1895)

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

f3035827

28 8月, 2019 2 次提交
- J
  Modify cast op and remove warning in argmax_test (#1894) · 5fe41d5c
  由 juncaipeng 提交于 8月 28, 2019
```
* modify cast op, test=develop

* modify cast op and remove warning in argmax_test, test=develop
```
  5fe41d5c
- H
  
  add floor op,elementwise_div op and assign op test=develop (#1882) · ca6974c5
  由 huzhiqiang 提交于 8月 28, 2019
  
  ca6974c5
26 8月, 2019 1 次提交

Add matmul op (#1837) · a1f4059f

由 Wilber 提交于 8月 26, 2019

* test=develop add matmul_op

* use lite::arm::math::sgemm func to implement matmul

* test=develop  pre-commit command to run clang-format

* Revert "test=develop  pre-commit command to run clang-format"

This reverts commit 3f56474f.

* test=develop pre-commit command to run clang-format

a1f4059f

25 8月, 2019 1 次提交
- Y
  
  leave tiny-publish out of third-party dependencies (#1853) · e5a76c98
  由 Yan Chunwei 提交于 8月 25, 2019
  
  e5a76c98