提交 · 3723451bdf104fe863c8c90a2903054ac64f21eb · PaddlePaddle / Paddle-Lite

23 12月, 2019 1 次提交
- Y
  
  [ARM] add grid_sampler op and ut, test=develop (#2598) · 3723451b
  由 yiicy 提交于 12月 23, 2019
  
  3723451b
20 12月, 2019 1 次提交
- W
  add var_conv_2d_relu pass test=develop (#2631) · 8304bc84
  由 Wilber 提交于 12月 20, 2019
```
add var_conv_2d + relu fuse pass
```
  8304bc84
19 12月, 2019 1 次提交
- W
  optimize cuda kernel test=develop (#2628) · 09aa15a5
  由 Wilber 提交于 12月 19, 2019
```
* optimize content-dnn cuda kernel
```
  09aa15a5
18 12月, 2019 1 次提交
- J
  Support Mask RCNN2 (#2588) · d1b7aec5
  由 juncaipeng 提交于 12月 18, 2019
```
* Support Mask RCNN2 (#2588)
```
  d1b7aec5
13 12月, 2019 1 次提交
- H
  [LITE][NPU][XPU] Refine subgraph pass, and support NPU/XPU model generation at... · d5434aa2
  由 hong19860320 提交于 12月 13, 2019
```
[LITE][NPU][XPU] Refine subgraph pass, and support NPU/XPU model generation at execution time (#2576)
```
  d5434aa2
10 12月, 2019 1 次提交
- Y
  
  [ARM] add instance norm op and ut, test=develop (#2578) · 9a3552db
  由 yiicy 提交于 12月 10, 2019
  
  9a3552db
08 12月, 2019 1 次提交
- L
  
  Add fc op on lite x86 platform (#2568) · d76c529a
  由 liu zhengxi 提交于 12月 08, 2019
  
  d76c529a
07 12月, 2019 1 次提交

由 juncaipeng 提交于 12月 07, 2019

* add arm split lod tensor, test=develop

* add arm merge lod tensor, test=develop

* update split merge lod tensor, test=develop

* add reduce_prob op, test=develop

* support mask_rcnn succeed, test=develop

c2f72cb3

27 11月, 2019 1 次提交
- fill_constant op support param shape can be tensor or tensorlist, test=develop (#2459) · 89df8f01
  由 myq406450149 提交于 11月 27, 2019
```
* fill_constant can support shape is tensor or tensorlist
```
  89df8f01
25 11月, 2019 1 次提交
- split op support param can be tensor or tensorlist,test=develop (#2474) · 493ea2ca
  由 myq406450149 提交于 11月 25, 2019
```
* split op upgrade
```
  493ea2ca
22 11月, 2019 2 次提交

update conv 2-pad to 4-pad (#2404) · 820eb6d4

由 HappyAngel 提交于 11月 22, 2019

* fix conv 2-pad to 4-pad

* fix compute conv shape

* fix pad, test=develop

* change conv_depthwise_3x3s1_fp.cc name to conv3x3s1p01_depthwise_fp32.cc to distinguish between conv3x3s1_depthwise_fp32.cc

* delete printf note in conv3x3s1, test=develop

* delete printf note, test=develop

* delete gem_sdot.h, test=develop

it is coped from __gemm_sdot_meta_.h

* update compute padding, test=develop

* fix padding size, must be 2 or 4. test=develop

* fix format in operators/conv_op.cc, test=develop

* change #if 0 to #if 1, test=develop

* put 2-pad to 4-pad in AttachImpl, test=develop

* fix clang-format error inn tests/math/connv_compute_test, test=develop

* fix x86 test result error, test=develop

* add asymmetric padding test case in liite/tests/math/conv_compute.cc, test=develop

* change paddings type to support dynamically modify, test=develop

* fix x86 build error in connv_compute_test, test=develop

* fix opencl build error, test=develop

* fix oopencl build error, test=develop

* fix  opencl/conv_compute build error, test=develop

* fix  opencl/conv_compute build error, test=develop

* fix format in kernels/opencl/conv_computte_ttest,test=develop

* fix build error, test=develop

fix build error in kernels/x86/conv_compute.h

820eb6d4

update pooling 2-padding to 4-padding (#2410) · a7f7d49b

由 HappyAngel 提交于 11月 22, 2019

* fix pooling bug and speed

* fix build error

* delete VLOGin pool, test=develop

* add openmp, test=develop

* fix lite/kernels/arm/pool_compute_test basic_pooling compute error bug, test=develop

* update pooling 2-pad to 4-pad, test=develop

* fix 2-pad to 4-pad in operators/pool_op.h, AttachKernel will set param, so 2-pad to 4-pad funcs should put in AttachKernel. test=ddevellop

* put 2-pad to 4-pad in AttachImpl, test=develop

* according to reviews, fix some format error. test=develop

* fix format errorr, add (). test=develop

* change paddings type to support dynamically modify, test=develop

* update padding type int other devices, test=develop

* fix x8d build error on shared_ptr, test=ddevelop

* fix formmat in operators pool_op.cc, test=develop

a7f7d49b

19 11月, 2019 3 次提交
- Y
  
  fix lrn param, align to fluid, test=develop (#2452) · 94255f6c
  由 yiicy 提交于 11月 19, 2019
  
  94255f6c
- H
  add x86 kernels: search_fc and sequence_topk_ave_pooling (#2443) · 68fe5b5c
  由 huzhiqiang 提交于 11月 18, 2019
```
* add x86 op and kernel : search_fc and sequence_topk_avg_pooling   for content-dnn model test=develop
```
  68fe5b5c
- Z
  [X86][CUDA] add attention_padding_mask op, x86 kernel, cuda kernel and unit tests (#2437) · ef6f7b84
  由 zhupengyang 提交于 11月 19, 2019
```
* [X86] add attention_padding_mask op, x86 kernel and unit test

test=develop

* [CUDA] add attention_padding_mask cuda kernel and unit test

test=develop
```
  ef6f7b84
18 11月, 2019 2 次提交
- P
  add search_group_padding op and x86 kernel, test=develop (#2440) · 1e88d1e8
  由 Pei Yang 提交于 11月 18, 2019
```
add search_group_padding op and x86 kernel
```
  1e88d1e8
- Z
  [X86][CUDA] add sequence_arithmetic op , x86 kernel, cuda kernel and unit test (#2436) · 8599c042
  由 zhupengyang 提交于 11月 18, 2019
```
* [X86][CUDA] add sequence_arithmetic op , x86 kernel, cuda kernel and unit test

test=develop

* add sequence_arithmetic cuda kernel unit test

test=develop
```
  8599c042
16 11月, 2019 1 次提交
- H
  
  [LITE][X86] Add search_aligned_mat_mul and search_seq_fc op for X86 (#2428) · 78f76834
  由 hong19860320 提交于 11月 16, 2019
  
  78f76834
15 11月, 2019 1 次提交

Add content-dnn ops (#2429) · 603b810f

由 juncaipeng 提交于 11月 15, 2019

* add search_seq_depadding x86 and cuda
* add match_matrix_tensor x86
* add search_grnn x86, no test

603b810f

14 11月, 2019 2 次提交
- W
  add var_conv_2d op, x86 kernel and unit test test=develop (#2422) · 9dcd9914
  由 Wilber 提交于 11月 14, 2019
```
- add var_conv_2d op

- add var_conv_2d x86 kernel

- add var_conv_2d x86 test
```
  9dcd9914
- L
  Fix the compile error for cuda and fix unit test (#2424) · 1f075a8b
  由 liu zhengxi 提交于 11月 14, 2019
```
* fix the compile error for cuda and fix unit test
```
  1f075a8b
13 11月, 2019 4 次提交
- W
  add sequence_reverse op and kerenl for arm and cuda test=develop (#2397) · acf09294
  由 Wilber 提交于 11月 13, 2019
```
- add sequence_reverse op

- add sequence_reverse kernel for x86 and cuda

- add sequence_reverse_test for x86 and cuda
```
  acf09294
- W
  add sequence_concat op kernel and test test=develop (#2414) · 8a1d942a
  由 Wilber 提交于 11月 13, 2019
```
- add sequence_concat op

- add sequence_concat kernel for x86 and cuda

- add sequence_concat_test for x86 and cuda
```
  8a1d942a
- L
  Update the ops to fluid (#2406) · 518a87ef
  由 liu zhengxi 提交于 11月 13, 2019
```
align the lite nearest， bilinear op to fluid on arm and cuda
```
  518a87ef
- J
  fix error for AxesTensorList in unsqueeze op, test=develop (#2411) · f4e06650
  由 juncaipeng 提交于 11月 13, 2019
```
* fix error for AxesTensorList in unsqueeze op
```
  f4e06650
12 11月, 2019 1 次提交
- J
  Upgrade concat and unsqueeze, test=develop (#2378) · 26470600
  由 juncaipeng 提交于 11月 12, 2019
```
* update concat and unsqueeze, test=develop
```
  26470600
06 11月, 2019 2 次提交

J
add channel_wise_dequantized_max_abs op and ChannelWiseDequantOpFuser (#2368) · ba85799c
由 juncaipeng 提交于 11月 06, 2019
```
* add channel_wise_dequantized_max_abs op and ChannelWiseDequantOpFuser, test=develop
```
ba85799c

update slice and reshape op and test on one op fake model test=develop (#2377) · e74609b7

由 Wilber 提交于 11月 06, 2019

update reshape op to support multiple input types of shape.
priority: input(ShapeTensor) > input(Shape) > attr(shape)

update slice op to support multiple iput types of starts and ends.
priority: input(StartsTensor) > input(StartsTensorList) > attr(starts)

e74609b7

05 11月, 2019 1 次提交
- L
  fix StepRNN model run related bugs (#2300) · 7f5c0ca1
  由 lijianshe02 提交于 11月 05, 2019
```
* fix step rnn model run bugs test=develop
```
  7f5c0ca1
28 10月, 2019 1 次提交

[LITE][XPU] initial support for XPU (#2202) · 06d058fe

由 hong19860320 提交于 10月 28, 2019

* Initial support for XPU
* Fix compiling errors of XPU
* Move XPU op kernel bridges from backends to kernels to fix deps order
* Change the namespace and directory of XPU bridges
* Add XPU SDK
* Fix header files and namespace of XPU SDK
* Add unit tests for relu and conv2d ops
* Restore the modification of paddle_api_test
* Supports simple model which contains only a relu layer
* Add compiling scripts for XPU
* Fix compiling errors of XPU
* Add comments for XPU LoadModel and BuildModel

06d058fe

23 10月, 2019 1 次提交

Enable pool2d, dropout, transpose and transpose2 op on x86 (#2226) · aa507f9b

由 liu zhengxi 提交于 10月 23, 2019

* enable pool2d op on x86 and add its unit tests, test=develop

* enable dropout op and add its unit tests, test=develop

* add tranpose, transpose2 op and add their unit tests, test=develop

aa507f9b

22 10月, 2019 2 次提交

Optimize quant_dequant (#2215) · f480d474

由 juncaipeng 提交于 10月 22, 2019

* Add DeleteQuantOpFuser
* Add fake_quantize_dequantize_moving_avg_abs_max_op
* Add DeleteQuantDequantOpFuser

f480d474

Transformer pr (#2214) · f0a6c1eb

由 TianXiaogang 提交于 10月 22, 2019

* feat: add beam_search_special function for support nlp model

* fix: add beam_search_compute kernel input and output

* feat: add assign op & copy_compute kernel

* feat: add fill_const_batch_size_like op & kernel

* feat: add layer_norm op and kernel and ut

* fix: fix some bugs
    fix mul_op infer_shape bug when x_dim_idx = 2, x_dims.size()=3 & y_dim_idx = 1, y_dims.size()=2
    fix elementwise_compute bug when y axis is all 1
    fix beam_search choose math_func wrong bug
    fix layer_norm get attr bug
    fix fill_constant_batch_size_like shape_set bug

* feat: add gather op and kernel & and transform ut

* feats: add ops and fix bugs to support transformer op
       fix type_cast passes to skip `while`
       fix elementwise infer_shape bug when x.dims=3 and y.dims={1} & axis=0
       fix lookup_table compute bug
       fix read_from_array/beam_search/increment/compate/gather ops data_type problems

* fix:
    transfomer ut add word read inferface
    fix copy/gather/norm/layer_norm include path problem

* fix:debug info

* fix: fix input reshape bug

* fix: fix norm bug

* style: style fix & test=develop

* style: fix operators cmakelist

* style: fix operators cmakelist; test=develop

* fix and test=develop

* fix and test=develop

* style: style fix; test=develop

f0a6c1eb

17 10月, 2019 1 次提交
- J
  
  add bilinear_interp_cuda_op, test=develop (#2197) · 4ac51a6b
  由 juncaipeng 提交于 10月 17, 2019
  
  4ac51a6b
15 10月, 2019 1 次提交

[NPU] Fix and refine the supporting of multi NPU models (#2037) · 7a731b7f

由 hong19860320 提交于 10月 15, 2019

* [NPU] Fix the bug of loading multi NPU models
test=develop

* [NPU] Use lite tensor to store NPU model, fix the management of multi NPU models, support loading NPU model from memory and reduce the modification of framework
test=develop

* [NPU] Remove redundant header files for NPU bridges,
test=develop

* [NPU] fix NPU deps
test=develop

* [NPU] refine the compiling script for NPU
test=develop

* [NPU] remove redundant subdirectory in lite/CMakeLists.txt
test=develop

* [NPU] Fix and refine NPU test case
test=develop

* [NPU] revoke the modification of other non-NPU modules
test=develop

* [NPU] Remove NPU bridges if target is tiny publish
test=develop

7a731b7f

14 10月, 2019 3 次提交
- J
  fix bug for reshape op, test=develop (#2141) · 421c6305
  由 juncaipeng 提交于 10月 14, 2019
```
* fix bug for reshape op, test=develop
```
  421c6305
- L
  fix asr modle related kernel bugs test=develop (#2179) · 792d898a
  由 lijianshe02 提交于 10月 14, 2019
```
* fix asr modle related kernel bugs test=develop
```
  792d898a
- J
  Optimize quant_dequant_fuse_pass (#2169) · 253acb80
  由 juncaipeng 提交于 10月 14, 2019
```
* optimize quant_dequant_fuse_pass, test=develop
```
  253acb80
11 10月, 2019 1 次提交

CUDA: can run yolov3 int8 (#2172) · 7931104f

由 Zhaolong Xing 提交于 10月 11, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

* can run yolov3 int8 test=develop

7931104f

18 9月, 2019 1 次提交

fix bias quantize error && fix clang build error (#2049) · 81dffbe8

由 Xiaoyang LI 提交于 9月 18, 2019

* fix gemm_int8, gemv-int8 and conv-int8 math function, add float bias

* change conv impl

* neon int8 kernel support float bias

* arm compute kernel support float bias

* add math_test target

* add tensor utils for testing, fix sgemm ut error

* add gemm_int8 unit test, support float bias

* fix build script

* add conv compute unit test for arm

* fix build script, test=develop

* fix fp32 dw conv3x3s1, test=develop

* add fp32 dw conv3x3s1, test=develop

* add armv7 fp32 dw conv3x3s1, test=develop

* add fp32 depthwise conv3x3s2, test=develop

* fix fp32 conv3x3 depthwise build error, test=develop

* fix gemm_like conv trans weights error, test=develop

* fix int8 depthwise conv3x3 error, test=develop

* turn on all test for arm fp32 conv, test=develop

* fix int8 conv1x1 error

* fix int8 direct conv3x3s1 error, test=develop

* fix int8 direct conv3x3s2, test=develop

* turn on all test for arm int8 conv, test=develop

* fix int8 fc error, change mobilenetv1-int8 ground-truth result to fluid, test=develop

* remove debug info, strip ut binary, test=develop

* fix conv compute error, test=develop

* change Init() to ReInitWhenNeeded(), test=develop

* fix code style, test=develop

* remote engine_test, test=develop

* fix building server tests error, test=develop

* fix sdot clang build error, test=develop

* fix sgemm ut timeout error, test=develop

* fix clang build error, test=develop

* turn off math basic test due to ci time out, test=develop

* fix conv_int8 ut error, test=develop

81dffbe8