提交 · 6b3c341f5a0a0a592cc9cd6e034b442349916f53 · PaddlePaddle / Paddle-Lite

22 11月, 2019 3 次提交

add NHWC NCHW transform, test=develop (#2381) · 6b3c341f

由 HappyAngel 提交于 11月 22, 2019

* add nhwc to nchw

* add layout in funcs

* change layout as extra, test=develop

* change make, test=develop

* use template class method to update layout NNCHHW and NHWC transform, test=develop

* fix cmake error, set layout to extra, test=develop

* fix test_layout_compute_arm test, its extra

* layout is extra, test=develop

* fix error in kernels/arm/layout_comput.cc when register kernel, DataLayout must be NCHW, test=develop

* delete extra note, test=develop

* delete extra test

* delete layout_test, test=develop

, its in tests/math/layout_comput_test

* delete extrat test, test=develop

6b3c341f

Y
[ARM] add sgemmc4 common and small kernel, support for winograd, test=develop (#2471) · 66d2ae25
由 yiicy 提交于 11月 22, 2019
```
* unfinish sgemmc4

* finish armv8 sgemmc4

* arm add sgemmc4 with deal with remain

* [ARM] add sgemmc4 small kernel, test=develop
```
66d2ae25

update pooling 2-padding to 4-padding (#2410) · a7f7d49b

由 HappyAngel 提交于 11月 22, 2019

* fix pooling bug and speed

* fix build error

* delete VLOGin pool, test=develop

* add openmp, test=develop

* fix lite/kernels/arm/pool_compute_test basic_pooling compute error bug, test=develop

* update pooling 2-pad to 4-pad, test=develop

* fix 2-pad to 4-pad in operators/pool_op.h, AttachKernel will set param, so 2-pad to 4-pad funcs should put in AttachKernel. test=ddevellop

* put 2-pad to 4-pad in AttachImpl, test=develop

* according to reviews, fix some format error. test=develop

* fix format errorr, add (). test=develop

* change paddings type to support dynamically modify, test=develop

* update padding type int other devices, test=develop

* fix x8d build error on shared_ptr, test=ddevelop

* fix formmat in operators pool_op.cc, test=develop

a7f7d49b

21 11月, 2019 1 次提交

石

fix cuda build error, test=develop (#2464) · d8ddbcc6

由石晓伟提交于 11月 21, 2019

* fix cuda building, test=develop

* remove sequence_pool from cmake because build error, test=develop

d8ddbcc6

20 11月, 2019 1 次提交
- Y
  [ARM] sgemv support transA, test=develop (#2453) · dde12f0d
  由 yiicy 提交于 11月 20, 2019
```
* [ARM] sgemv support transA, test=develop

* add sgemv ut, test=develop
```
  dde12f0d
19 11月, 2019 2 次提交
- H
  
  [LITE][CUDA] Add CUDA kernel for search_aligned_mat_mul and search_seq_fc Op (#2449) · 8373aec5
  由 hong19860320 提交于 11月 19, 2019
  
  8373aec5
- H
  add x86 kernels: search_fc and sequence_topk_ave_pooling (#2443) · 68fe5b5c
  由 huzhiqiang 提交于 11月 18, 2019
```
* add x86 op and kernel : search_fc and sequence_topk_avg_pooling   for content-dnn model test=develop
```
  68fe5b5c
18 11月, 2019 1 次提交

[LITE][OPENCL] Enable full and light api for OpenCL (#2331) · d242bdfb

由 Yuan Shuai 提交于 11月 18, 2019

* Fix bug target for kHost and kARM not equal. test=develop

* Fix license. test=develop

* add debug -g option. test=develop

* enable opencl demo. test=develop

* Fix model_optimize_tool found no opencl kernel. test=develop

* add more vlog. test=develop

* remove macro LITE_WITH_OPENCL, LITE_WITH_FPGA in passes. test=develop

* Fix valid_places in mobilenetv1_test. test=develop

* Fix bug of find no real output of fetch, after tool OPs of optimzer passes. test=develop

* Fix vlog as log message in model_optimize_tool. test=develop

* fix miscs. test=develop

* fix comment. test=develop

* Fix misspell of opencl, fpga kernels name in lite/api/CMakeLists.txt. test=develop

* add opencl macro in full_api of demo. test=develop

d242bdfb

17 11月, 2019 1 次提交
- J
  Add cuda match_matrix_tensor op and test (#2434) · ce21ff5d
  由 juncaipeng 提交于 11月 17, 2019
```
* add cuda match_matrix_tensor op and test, test=develop
```
  ce21ff5d
14 11月, 2019 1 次提交
- H
  
  [LITE][NPU] Upgrade HiAI DDK from 300 to 310 (#2423) · c1837d76
  由 hong19860320 提交于 11月 14, 2019
  
  c1837d76
13 11月, 2019 1 次提交
- L
  Update the ops to fluid (#2406) · 518a87ef
  由 liu zhengxi 提交于 11月 13, 2019
```
align the lite nearest， bilinear op to fluid on arm and cuda
```
  518a87ef
11 11月, 2019 1 次提交

fix pool bug and speed, test=develop (#2385) · d197de00

由 HappyAngel 提交于 11月 11, 2019

* fix pooling bug and speed

* fix build error

* delete VLOG in pool, test=develop

* add openmp, test=develop

* fix lite/kernels/arm/pool_compute_test basic_pooling compute error bug, test=develop

d197de00

08 11月, 2019 1 次提交

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the... · aacba6f5

由 hong19860320 提交于 11月 08, 2019

[LITE][NPU] Add fusion_elementwise_add_activation bridge, remove the limitation of the dimensions of input tensors in the graph compute kernel, and refine log message (#2395)

test=develop

aacba6f5

07 11月, 2019 1 次提交

fix jit_matmul bug according to paddle pr#20948 test=develop (#2392) · f51c4891

由 Wilber 提交于 11月 07, 2019

fix jit::matmul bug. Input x shape is (m, k), weight shape is (k, n). When k < 512, m==1, and n is a multiple of 16, the weight pointer is not correctly updated in the group calculation in the implementation of jit::matmul, resulting in the result diff

f51c4891

30 10月, 2019 1 次提交
- H
  
  [LITE][NPU] Use FullConnection op to solve the compatibility between Kirin 810 and 990 (#2283) · f3e5d3e5
  由 hong19860320 提交于 10月 30, 2019
  
  f3e5d3e5
29 10月, 2019 2 次提交

Add tanh op and gelu op for x86 platform (#2265) · 605a309c

由 liu zhengxi 提交于 10月 29, 2019

* add tanh op in x86 platform and its unittest, test=develop

* add gelu op on x86 platform and add its unittests, test=develop

* update depends for math_function for activation for gelu, test=develop

605a309c

H
[LITE][NPU] Add supporting for Huawei offical DDK (#2262) · dc2b853e
由 hong19860320 提交于 10月 29, 2019
```
* Add supporting for Huawei offical DDK
* Fix the param of graph op in NPU graph computing kernel
```
dc2b853e

28 10月, 2019 1 次提交

[LITE][XPU] initial support for XPU (#2202) · 06d058fe

由 hong19860320 提交于 10月 28, 2019

* Initial support for XPU
* Fix compiling errors of XPU
* Move XPU op kernel bridges from backends to kernels to fix deps order
* Change the namespace and directory of XPU bridges
* Add XPU SDK
* Fix header files and namespace of XPU SDK
* Add unit tests for relu and conv2d ops
* Restore the modification of paddle_api_test
* Supports simple model which contains only a relu layer
* Add compiling scripts for XPU
* Fix compiling errors of XPU
* Add comments for XPU LoadModel and BuildModel

06d058fe

24 10月, 2019 1 次提交

Fix gemv int8 error (#2249) · 3406a11a

由 Xiaoyang LI 提交于 10月 24, 2019

* remove log in reshape, fix conv error when padding size=4, test=develop

* fix style, test=develop

* remove useless code, test=develop

* remove redundant model test file, test=develop

* change cluster to power_mode, test=develop

* fix build error, test=develop

* change cluster to power_mode, test=develop

* change opt_nb to use_optimize_nb, test=develop

* null, test=develop

* add gemv-int8 test, fix clang build error, test=develop

* fix gemv-int8 error when build with clang, test=develop

3406a11a

23 10月, 2019 3 次提交
- W
  modify yolobox_cuda to support multiple runs (#2245) · a4a19ba4
  由 Wilber 提交于 10月 23, 2019
```
* modify yolobox_cuda to support multiple runs test=develop
```
  a4a19ba4
- L
  Enable pool2d, dropout, transpose and transpose2 op on x86 (#2226) · aa507f9b
  由 liu zhengxi 提交于 10月 23, 2019
```
* enable pool2d op on x86 and add its unit tests, test=develop

* enable dropout op and add its unit tests, test=develop

* add tranpose, transpose2 op and add their unit tests, test=develop
```
  aa507f9b
- S
  add python api (#2225) · 76e74ef1
  由 sangoly 提交于 10月 23, 2019
```
* [python api] init add python api test=develop
```
  76e74ef1
22 10月, 2019 1 次提交

Transformer pr (#2214) · f0a6c1eb

由 TianXiaogang 提交于 10月 22, 2019

* feat: add beam_search_special function for support nlp model

* fix: add beam_search_compute kernel input and output

* feat: add assign op & copy_compute kernel

* feat: add fill_const_batch_size_like op & kernel

* feat: add layer_norm op and kernel and ut

* fix: fix some bugs
    fix mul_op infer_shape bug when x_dim_idx = 2, x_dims.size()=3 & y_dim_idx = 1, y_dims.size()=2
    fix elementwise_compute bug when y axis is all 1
    fix beam_search choose math_func wrong bug
    fix layer_norm get attr bug
    fix fill_constant_batch_size_like shape_set bug

* feat: add gather op and kernel & and transform ut

* feats: add ops and fix bugs to support transformer op
       fix type_cast passes to skip `while`
       fix elementwise infer_shape bug when x.dims=3 and y.dims={1} & axis=0
       fix lookup_table compute bug
       fix read_from_array/beam_search/increment/compate/gather ops data_type problems

* fix:
    transfomer ut add word read inferface
    fix copy/gather/norm/layer_norm include path problem

* fix:debug info

* fix: fix input reshape bug

* fix: fix norm bug

* style: style fix & test=develop

* style: fix operators cmakelist

* style: fix operators cmakelist; test=develop

* fix and test=develop

* fix and test=develop

* style: style fix; test=develop

f0a6c1eb

21 10月, 2019 2 次提交
- 石
  link static library with cuda, test=develop (#2228) · 7c69b6b4
  由石晓伟提交于 10月 21, 2019
```
* add static libraries of cuda, test=develop

* update cuda make
```
  7c69b6b4
- to support yolov3 unet alexnet can run on tx2 (#2216) · 57d8e42e
  由 myq406450149 提交于 10月 21, 2019
```
* add gpu kernel mul pool relu scale softmax dropout bilinear_interp and can run in tx2

* rm GREATER_EQUAL
```
  57d8e42e
17 10月, 2019 2 次提交

fix npu path (#2210) · 7c722a37

由 zhupengyang 提交于 10月 17, 2019

* move lite/backends/npu/bridges --> lite/kernels/npu/

test=develop

* fix namespace for npu

test=develop

* mv npu runtime file to lite/backends/npu

test=develop

7c722a37

speedup fp32 depthwise conv · 2f6d5f9e

由 HappyAngel 提交于 10月 17, 2019

* update con_dw

* update

* add conv_depthwise_3x3s1.cc and conv_depthwise_3x3s2.cc

* add conv_depthwise_3x3s1_fp32 and conv_depthwise_3x3s2_fp32

* add new conv_dw

* only support conv_dw pad=0, 1

* add conv_dw_s1 conv_dw_s2 fp32

*     //conv2_func _impl2{nullptr};
update conv_dw, add conv_3x3s1 and conv_3x3s2, pad=[0,1]

* fix format, test=develop

* fix formmat, test=develop

2f6d5f9e

16 10月, 2019 1 次提交
- Z
  Ban feed and fetch op during inference (#2198) · 75e8a6fc
  由 Zhaolong Xing 提交于 10月 16, 2019
```
* init: delete feed and fetch op, using zero copy
test=develop

* delete the unused test
test=develop
```
  75e8a6fc
15 10月, 2019 2 次提交

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of... · 72c11758

由 Yuan Shuai 提交于 10月 15, 2019

[LITE][OPENCL] Fix layout, target pass for OpenCL, add macro of CONVERT_TYPE_TO and READ/WRITE image, memory reuse in ResetLazyImage2D (#2170)

* add macro of CONVERT_TYPE_TO and READ/WRITE image. test=develop

* add data type control. test=develop

* fix io op as general layout and precision. test=develop

* Fix memory reuse strategy for opencl image2d. test=develop

* remove std::array, std::map in about opencl backend. test=develop

72c11758

[NPU] Fix and refine the supporting of multi NPU models (#2037) · 7a731b7f

由 hong19860320 提交于 10月 15, 2019

* [NPU] Fix the bug of loading multi NPU models
test=develop

* [NPU] Use lite tensor to store NPU model, fix the management of multi NPU models, support loading NPU model from memory and reduce the modification of framework
test=develop

* [NPU] Remove redundant header files for NPU bridges,
test=develop

* [NPU] fix NPU deps
test=develop

* [NPU] refine the compiling script for NPU
test=develop

* [NPU] remove redundant subdirectory in lite/CMakeLists.txt
test=develop

* [NPU] Fix and refine NPU test case
test=develop

* [NPU] revoke the modification of other non-NPU modules
test=develop

* [NPU] Remove NPU bridges if target is tiny publish
test=develop

7a731b7f

14 10月, 2019 2 次提交
- Z
  align yolov3 cuda int8 (#2183) · 80d35725
  由 Zhaolong Xing 提交于 10月 14, 2019
```
test=develop
```
  80d35725
- L
  fix asr modle related kernel bugs test=develop (#2179) · 792d898a
  由 lijianshe02 提交于 10月 14, 2019
```
* fix asr modle related kernel bugs test=develop
```
  792d898a
11 10月, 2019 3 次提交

J

add rsqrt op, test=develop (#2176) · dfce4621
由 juncaipeng 提交于 10月 11, 2019

dfce4621

CUDA: can run yolov3 int8 (#2172) · 7931104f

由 Zhaolong Xing 提交于 10月 11, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

* can run yolov3 int8 test=develop

7931104f

[LITE][OPENCL] support image2d type (#2158) · 77cdbdce

由 Yuan Shuai 提交于 10月 11, 2019

* [LITE][OPENCL] support image2d. test=develop

* add context changed with consider image*. test=develop

* add layout, relu image kernels. test=develop

* replace image_data with data, mutable_image_data with mutable_data, test=develop

* comment unused var. test=develop

* remove unused var. test=develop

77cdbdce

10 10月, 2019 1 次提交
- W
  fix yolobox_cuda bug · f4ac2768
  由 Wilber 提交于 10月 10, 2019
```
* fix yolobox_cuda bug 
* update code format
```
  f4ac2768
09 10月, 2019 1 次提交

improve dw conv performance · 4b9df8fb

由 yiicy 提交于 10月 09, 2019

*  imporve prepack_input func speed in int8 3x3s1 dw conv

* fix code style

* fix code style

* improve 3x3s1 dw fp32 conv speed a little

* arm add 5x5s1 int8 dw conv, test=develop

4b9df8fb

27 9月, 2019 1 次提交

can run yolov3 fp32 on cuda devices (#2092) · 3d6d744f

由 Zhaolong Xing 提交于 9月 27, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

3d6d744f

25 9月, 2019 1 次提交
- X
  
  add workspace compute funcs for direct conv, test=develop (#2132) · 61647c35
  由 Xiaoyang LI 提交于 9月 25, 2019
  
  61647c35
19 9月, 2019 1 次提交

石

add full_api_static target and fix building errors, test=develop (#2064) · eef7ea0f

由石晓伟提交于 9月 19, 2019

* add full_api_static target and fix building errors, test=develop

* fix build errors, test=develop

* fix code style, test=develop

* fix lite/model_parser/pb/var_desc.cc, test=develop

* fix building errors, test=develop

* modify lite/tools/debug/CMakeLists.txt, test=develop

eef7ea0f