提交 · 3d5e261be64191a3583cb3007763ce1777ef3ae1 · PaddlePaddle / Paddle-Lite

20 11月, 2019 1 次提交
- Y
  [ARM] sgemv support transA, test=develop (#2453) · 3d5e261b
  由 yiicy 提交于 11月 20, 2019
```
* [ARM] sgemv support transA, test=develop

* add sgemv ut, test=develop
```
  3d5e261b
19 11月, 2019 1 次提交
- Y
  
  fix lrn param, align to fluid, test=develop (#2452) · bcad1580
  由 yiicy 提交于 11月 19, 2019
  
  bcad1580
14 11月, 2019 1 次提交
- P
  Update lookup_table op on arm x86，add lookup_table_v2_op (#2405) · 24029728
  由 Pei Yang 提交于 11月 14, 2019
```
* update lookup_table arm x86, test=develop

* add lookup_table_v2_op for compatibility, test=develop
```
  24029728
13 11月, 2019 1 次提交
- L
  Update the ops to fluid (#2406) · c77e4074
  由 liu zhengxi 提交于 11月 13, 2019
```
align the lite nearest， bilinear op to fluid on arm and cuda
```
  c77e4074
12 11月, 2019 1 次提交
- J
  Upgrade concat and unsqueeze, test=develop (#2378) · 284e8166
  由 juncaipeng 提交于 11月 12, 2019
```
* update concat and unsqueeze, test=develop
```
  284e8166
11 11月, 2019 1 次提交

fix pool bug and speed, test=develop (#2385) · f85c3689

由 HappyAngel 提交于 11月 11, 2019

* fix pooling bug and speed

* fix build error

* delete VLOG in pool, test=develop

* add openmp, test=develop

* fix lite/kernels/arm/pool_compute_test basic_pooling compute error bug, test=develop

f85c3689

08 11月, 2019 1 次提交

Move muliti class kernel back to basic (#2396) · 7ea34b1b

由 huzhiqiang 提交于 11月 07, 2019

* move multiclass_nms kernel back to host test=develop

* move layer_norm OP and arm_kernel into extra type since it's added after release/v2.0-beta1 and not related with CV test=develop

* fix code_style test=develop

7ea34b1b

07 11月, 2019 2 次提交

T
fix:fix deviceinfo worksapce to tls · 30300c98
由 TianXiaogang 提交于 11月 07, 2019
```
mv deviceinfo.workspace and other relative member to thread_local_storage
```
30300c98

check arm kernels type to make sure all_library_links work normally (#2386) · 6b38eab8

由 huzhiqiang 提交于 11月 06, 2019

We have changed 11 arm_kernels into extra type in #2347 , which has caused test_compiling failure. In this PR , we move their 11 related arm_kernel_test into build_extra=ON

6b38eab8

06 11月, 2019 3 次提交

update slice and reshape op and test on one op fake model test=develop (#2377) · 126964a1

由 Wilber 提交于 11月 06, 2019

update reshape op to support multiple input types of shape.
priority: input(ShapeTensor) > input(Shape) > attr(shape)

update slice op to support multiple iput types of starts and ends.
priority: input(StartsTensor) > input(StartsTensorList) > attr(starts)

126964a1

fix fill_constant kernel bug test=develop (#2376) · f607870c

由 Wilber 提交于 11月 06, 2019

fill_constant kernel only registered float type, only the float data type is produced, which is obviously a bug.

Now, produce data based on the data type attr.

By the way, fix the cast kernel bug.

f607870c

H

change arm conv_kernels into basic to support mobilenetv1 test=develop (#2375) · 62bbf501
由 huzhiqiang 提交于 11月 06, 2019

62bbf501

04 11月, 2019 1 次提交
- H
  Move new op kernel into extra (#2348) · a438d0dc
  由 huzhiqiang 提交于 11月 04, 2019
```
* move some basic ops into extra type to reduce library size test=develop (#2347)
```
  a438d0dc
01 11月, 2019 1 次提交
- T
  
  fix: fix conv_direct && test=develop (#2314) · ecca7325
  由 TianXiaogang 提交于 11月 01, 2019
  
  ecca7325
29 10月, 2019 1 次提交
- G
  modify layer_norm_compute_test.cc (#2261) · f2b9ca34
  由 guofei 提交于 10月 29, 2019
```
test=develop
```
  f2b9ca34
23 10月, 2019 1 次提交
- T
  fix: fix fc batch bug (#2244) · ea4a5854
  由 TianXiaogang 提交于 10月 23, 2019
```
* fix: fix fc batch bug

* fix:fix fc bug;test=develop
```
  ea4a5854
22 10月, 2019 1 次提交

Transformer pr (#2214) · 330644b0

由 TianXiaogang 提交于 10月 22, 2019

* feat: add beam_search_special function for support nlp model

* fix: add beam_search_compute kernel input and output

* feat: add assign op & copy_compute kernel

* feat: add fill_const_batch_size_like op & kernel

* feat: add layer_norm op and kernel and ut

* fix: fix some bugs
    fix mul_op infer_shape bug when x_dim_idx = 2, x_dims.size()=3 & y_dim_idx = 1, y_dims.size()=2
    fix elementwise_compute bug when y axis is all 1
    fix beam_search choose math_func wrong bug
    fix layer_norm get attr bug
    fix fill_constant_batch_size_like shape_set bug

* feat: add gather op and kernel & and transform ut

* feats: add ops and fix bugs to support transformer op
       fix type_cast passes to skip `while`
       fix elementwise infer_shape bug when x.dims=3 and y.dims={1} & axis=0
       fix lookup_table compute bug
       fix read_from_array/beam_search/increment/compate/gather ops data_type problems

* fix:
    transfomer ut add word read inferface
    fix copy/gather/norm/layer_norm include path problem

* fix:debug info

* fix: fix input reshape bug

* fix: fix norm bug

* style: style fix & test=develop

* style: fix operators cmakelist

* style: fix operators cmakelist; test=develop

* fix and test=develop

* fix and test=develop

* style: style fix; test=develop

330644b0

21 10月, 2019 1 次提交

add cuda op(pool & softmax), support conv with padding_algorithm · 1a18d682

由 yiicy 提交于 10月 21, 2019

* cuda add softmax and pool op

* * fix armlinux can find sys/system_properties.h
* conv add padding_algorithm
test=develop

* delete padding_algorithm in op param, test=develop

* fix bugs, test=develop

1a18d682

17 10月, 2019 1 次提交

speedup fp32 depthwise conv · c5cd78ab

由 HappyAngel 提交于 10月 17, 2019

* update con_dw

* update

* add conv_depthwise_3x3s1.cc and conv_depthwise_3x3s2.cc

* add conv_depthwise_3x3s1_fp32 and conv_depthwise_3x3s2_fp32

* add new conv_dw

* only support conv_dw pad=0, 1

* add conv_dw_s1 conv_dw_s2 fp32

*     //conv2_func _impl2{nullptr};
update conv_dw, add conv_3x3s1 and conv_3x3s2, pad=[0,1]

* fix format, test=develop

* fix formmat, test=develop

c5cd78ab

12 10月, 2019 1 次提交

fix conv_transpose error (#2165) · 9a464d63

由 Xiaoyang LI 提交于 10月 12, 2019

* fix conv_transpose error

* fix build error, enable basic test of conv_transpose, test=develop

9a464d63

11 10月, 2019 1 次提交
- J
  
  add rsqrt op, test=develop (#2176) · 78ddd64d
  由 juncaipeng 提交于 10月 11, 2019
  
  78ddd64d
09 10月, 2019 1 次提交

improve dw conv performance · 498a30cf

由 yiicy 提交于 10月 09, 2019

*  imporve prepack_input func speed in int8 3x3s1 dw conv

* fix code style

* fix code style

* improve 3x3s1 dw fp32 conv speed a little

* arm add 5x5s1 int8 dw conv, test=develop

498a30cf

25 9月, 2019 1 次提交
- X
  
  add workspace compute funcs for direct conv, test=develop (#2132) · f03217b4
  由 Xiaoyang LI 提交于 9月 25, 2019
  
  f03217b4
23 9月, 2019 1 次提交
- J
  add cast from uint8 to float, test=develop (#2080) · babde352
  由 juncaipeng 提交于 9月 23, 2019
```
* add cast from uint8 to float, test=develop
```
  babde352
18 9月, 2019 1 次提交

fix bias quantize error && fix clang build error (#2049) · 8d6f475e

由 Xiaoyang LI 提交于 9月 18, 2019

* fix gemm_int8, gemv-int8 and conv-int8 math function, add float bias

* change conv impl

* neon int8 kernel support float bias

* arm compute kernel support float bias

* add math_test target

* add tensor utils for testing, fix sgemm ut error

* add gemm_int8 unit test, support float bias

* fix build script

* add conv compute unit test for arm

* fix build script, test=develop

* fix fp32 dw conv3x3s1, test=develop

* add fp32 dw conv3x3s1, test=develop

* add armv7 fp32 dw conv3x3s1, test=develop

* add fp32 depthwise conv3x3s2, test=develop

* fix fp32 conv3x3 depthwise build error, test=develop

* fix gemm_like conv trans weights error, test=develop

* fix int8 depthwise conv3x3 error, test=develop

* turn on all test for arm fp32 conv, test=develop

* fix int8 conv1x1 error

* fix int8 direct conv3x3s1 error, test=develop

* fix int8 direct conv3x3s2, test=develop

* turn on all test for arm int8 conv, test=develop

* fix int8 fc error, change mobilenetv1-int8 ground-truth result to fluid, test=develop

* remove debug info, strip ut binary, test=develop

* fix conv compute error, test=develop

* change Init() to ReInitWhenNeeded(), test=develop

* fix code style, test=develop

* remote engine_test, test=develop

* fix building server tests error, test=develop

* fix sdot clang build error, test=develop

* fix sgemm ut timeout error, test=develop

* fix clang build error, test=develop

* turn off math basic test due to ci time out, test=develop

* fix conv_int8 ut error, test=develop

8d6f475e

17 9月, 2019 1 次提交
- W
  modify norm kernel to run caffe_facedetection model (#2008) · 133a40d2
  由 Wilber 提交于 9月 17, 2019
```
* modify norm kernel to run caffe_facedetection model

* reserve bind norm, remove calc of norm output
```
  133a40d2
12 9月, 2019 3 次提交
- G
  
  enable native compiling on raspberry pi and rk3399 (#2021) · 9fae3170
  由 guofei 提交于 9月 12, 2019
  
  9fae3170
- W
  add unsqueeze and range op (x2paddle) (#1988) · cca0aec6
  由 Wilber 提交于 9月 12, 2019
```
* add unsqueeze and range op. modify concat op test=develop

* modify exception in range_test_x86
```
  cca0aec6
- W
  
  add min_max_aspect_ratios_order attr in prior box op test=develop (#2016) · f04ed39c
  由 Wilber 提交于 9月 12, 2019
  
  f04ed39c
11 9月, 2019 1 次提交
- Y
  
  make model_optimize_tool run on host (#1990) · d72dc4d2
  由 Yan Chunwei 提交于 9月 11, 2019
  
  d72dc4d2
10 9月, 2019 1 次提交
- W
  
  add elementwise_sub and modify argmax (#1964) · 192320c4
  由 Wilber 提交于 9月 10, 2019
  
  192320c4
09 9月, 2019 1 次提交
- J
  add assign_value and hard_sigmoid, add fluid_type (#1983) · 9796c57d
  由 juncaipeng 提交于 9月 09, 2019
```
* add assign_value op, arm kernel and test, add fluid_type, test=develop

* add hard_sigmoid, test=develop
```
  9796c57d
03 9月, 2019 1 次提交
- H
  
  create backends directory and move hardware backends into it (#1954) · fede4a1c
  由 huzhiqiang 提交于 9月 03, 2019
  
  fede4a1c
02 9月, 2019 2 次提交

Add ops and fix bugs for Faster RCNN (#1942) · cfd5abe5

由 juncaipeng 提交于 9月 02, 2019

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

* add box_clip op and fix tensor slice bug

* remove add four op twice

* rewrite the implement for box_coder and sequence_expand, add faster_rcnn_test, test=develop

* fix test bug of box_clip in x86 server, test=develop

cfd5abe5

fix elementwise_op bug when the shape of input y is 1, test=develop (#1924) · 69217a24

由 juncaipeng 提交于 9月 02, 2019

* fix elementwise_op bug when the shape of input y is 1, test=develop

* fix elementwise ops bug when the shape of input y is 1, test=develop

69217a24

30 8月, 2019 1 次提交

Modify the slice op, stack op, reduce_mean op (#1921) · 440aae7a

由 liu zhengxi 提交于 8月 30, 2019

* add stack op and add reduce_mean op and their unit tests

* modify stack op output name and modify the for loop in reduce_mean op

* add HasAttr for slice op

440aae7a

29 8月, 2019 3 次提交

Add yolo_box_cuda multiclass_nms_host kernel. (#1908) · 5752dbd7

由 Wilber 提交于 8月 29, 2019

* add yolo_box_compute cuda

* move multiclass_nms(arm) to host

* add lod in scale op

* add yolo_box_cuda cmake config

* modify shuffle_channel_fuse and transpose_softmax_transpose_fuse to support run ssd model. test=develop

* reshape and transpose op don't have xshape output.

* modify yolo_box_compute_cuda, use tensor to manage cuda memory test=develop

* add yolo_box use kernel test=develop

5752dbd7

L

add stack op and add reduce_mean op and their unit tests (#1888) · 8ccd01a6
由 liu zhengxi 提交于 8月 29, 2019

8ccd01a6

ad ops for faster rcnn, including affine_channel, anchor_generator,... · f3035827

由 juncaipeng 提交于 8月 29, 2019

ad ops for faster rcnn, including affine_channel, anchor_generator, generate_proposals and roi_align (#1895)

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

f3035827

28 8月, 2019 1 次提交
- J
  Modify cast op and remove warning in argmax_test (#1894) · 5fe41d5c
  由 juncaipeng 提交于 8月 28, 2019
```
* modify cast op, test=develop

* modify cast op and remove warning in argmax_test, test=develop
```
  5fe41d5c