提交 · 7931104f5f5211e4b8e370481e1f8ff04b7a57ac · PaddlePaddle / Paddle-Lite

11 10月, 2019 2 次提交

CUDA: can run yolov3 int8 (#2172) · 7931104f

由 Zhaolong Xing 提交于 10月 11, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

* can run yolov3 int8 test=develop

7931104f

[LITE][OPENCL] support image2d type (#2158) · 77cdbdce

由 Yuan Shuai 提交于 10月 11, 2019

* [LITE][OPENCL] support image2d. test=develop

* add context changed with consider image*. test=develop

* add layout, relu image kernels. test=develop

* replace image_data with data, mutable_image_data with mutable_data, test=develop

* comment unused var. test=develop

* remove unused var. test=develop

77cdbdce

10 10月, 2019 1 次提交
- W
  fix yolobox_cuda bug · f4ac2768
  由 Wilber 提交于 10月 10, 2019
```
* fix yolobox_cuda bug 
* update code format
```
  f4ac2768
09 10月, 2019 1 次提交

improve dw conv performance · 4b9df8fb

由 yiicy 提交于 10月 09, 2019

*  imporve prepack_input func speed in int8 3x3s1 dw conv

* fix code style

* fix code style

* improve 3x3s1 dw fp32 conv speed a little

* arm add 5x5s1 int8 dw conv, test=develop

4b9df8fb

27 9月, 2019 1 次提交

can run yolov3 fp32 on cuda devices (#2092) · 3d6d744f

由 Zhaolong Xing 提交于 9月 27, 2019

* add conv int8 support(in condition which the input or output channel not be the times of 4)
add add_kernel for cuda.

* can run yolov3 fp32
test=develop

* 1. fix bug with yolov3 run
test=develop

3d6d744f

25 9月, 2019 1 次提交
- X
  
  add workspace compute funcs for direct conv, test=develop (#2132) · 61647c35
  由 Xiaoyang LI 提交于 9月 25, 2019
  
  61647c35
23 9月, 2019 1 次提交
- J
  add cast from uint8 to float, test=develop (#2080) · 9941d746
  由 juncaipeng 提交于 9月 23, 2019
```
* add cast from uint8 to float, test=develop
```
  9941d746
20 9月, 2019 1 次提交
- P
  
  refine concat cuda kernel, test=develop (#2081) · cef884e5
  由 Pei Yang 提交于 9月 20, 2019
  
  cef884e5
19 9月, 2019 1 次提交

石

add full_api_static target and fix building errors, test=develop (#2064) · eef7ea0f

由石晓伟提交于 9月 19, 2019

* add full_api_static target and fix building errors, test=develop

* fix build errors, test=develop

* fix code style, test=develop

* fix lite/model_parser/pb/var_desc.cc, test=develop

* fix building errors, test=develop

* modify lite/tools/debug/CMakeLists.txt, test=develop

eef7ea0f

18 9月, 2019 1 次提交

fix bias quantize error && fix clang build error (#2049) · 81dffbe8

由 Xiaoyang LI 提交于 9月 18, 2019

* fix gemm_int8, gemv-int8 and conv-int8 math function, add float bias

* change conv impl

* neon int8 kernel support float bias

* arm compute kernel support float bias

* add math_test target

* add tensor utils for testing, fix sgemm ut error

* add gemm_int8 unit test, support float bias

* fix build script

* add conv compute unit test for arm

* fix build script, test=develop

* fix fp32 dw conv3x3s1, test=develop

* add fp32 dw conv3x3s1, test=develop

* add armv7 fp32 dw conv3x3s1, test=develop

* add fp32 depthwise conv3x3s2, test=develop

* fix fp32 conv3x3 depthwise build error, test=develop

* fix gemm_like conv trans weights error, test=develop

* fix int8 depthwise conv3x3 error, test=develop

* turn on all test for arm fp32 conv, test=develop

* fix int8 conv1x1 error

* fix int8 direct conv3x3s1 error, test=develop

* fix int8 direct conv3x3s2, test=develop

* turn on all test for arm int8 conv, test=develop

* fix int8 fc error, change mobilenetv1-int8 ground-truth result to fluid, test=develop

* remove debug info, strip ut binary, test=develop

* fix conv compute error, test=develop

* change Init() to ReInitWhenNeeded(), test=develop

* fix code style, test=develop

* remote engine_test, test=develop

* fix building server tests error, test=develop

* fix sdot clang build error, test=develop

* fix sgemm ut timeout error, test=develop

* fix clang build error, test=develop

* turn off math basic test due to ci time out, test=develop

* fix conv_int8 ut error, test=develop

81dffbe8

17 9月, 2019 2 次提交
- W
  modify norm kernel to run caffe_facedetection model (#2008) · 71bb3188
  由 Wilber 提交于 9月 17, 2019
```
* modify norm kernel to run caffe_facedetection model

* reserve bind norm, remove calc of norm output
```
  71bb3188
- L
  add fill_constant_batch_size_like op and add its unittest (#2044) · f5338469
  由 liu zhengxi 提交于 9月 17, 2019
```
* add fill_constant_batch_size_like op and add its unittest
```
  f5338469
16 9月, 2019 1 次提交
- L
  Gru op (#2002) · eb42f9ee
  由 lhl960107 提交于 9月 16, 2019
```
* add x86 gru&&relu&&sequence_expand_as op test=develop
```
  eb42f9ee
12 9月, 2019 6 次提交
- L
  
  add matmul op kernels for asr test=develop (#2032) · b6379fed
  由 lijianshe02 提交于 9月 12, 2019
  
  b6379fed
- G
  
  enable native compiling on raspberry pi and rk3399 (#2021) · e0b4b5c9
  由 guofei 提交于 9月 12, 2019
  
  e0b4b5c9
- W
  add unsqueeze and range op (x2paddle) (#1988) · 3c08f676
  由 Wilber 提交于 9月 12, 2019
```
* add unsqueeze and range op. modify concat op test=develop

* modify exception in range_test_x86
```
  3c08f676
- W
  
  add min_max_aspect_ratios_order attr in prior box op test=develop (#2016) · cf84d42b
  由 Wilber 提交于 9月 12, 2019
  
  cf84d42b
- L
  
  add elementwise op function and add elementwise add/sub kernels test=develop (#2020) · 806ba6e7
  由 lijianshe02 提交于 9月 12, 2019
  
  806ba6e7
- W
  add transpose kernel for cuda test=develop (#1997) · cba5736f
  由 Wilber 提交于 9月 12, 2019
```
add transpose kernel for cuda
```
  cba5736f
11 9月, 2019 2 次提交
- Y
  
  make model_optimize_tool run on host (#1990) · 83d4b0e8
  由 Yan Chunwei 提交于 9月 11, 2019
  
  83d4b0e8
- L
  add slice op, reshape op, reshape2 op, squeeze op, squeeze2 op for x86 (#2005) · 13bbd2b8
  由 liu zhengxi 提交于 9月 11, 2019
```
add slice op, reshape op,  reshape2 op, squeeze op and squeeze2 op and their unittests for x86
```
  13bbd2b8
10 9月, 2019 3 次提交
- L
  
  add x86 softmax kernel and fix jit compute bugs test=develop (#2007) · 81132a32
  由 lijianshe02 提交于 9月 10, 2019
  
  81132a32
- W
  
  add elementwise_sub and modify argmax (#1964) · 62ea82d0
  由 Wilber 提交于 9月 10, 2019
  
  62ea82d0
- T
  
  fix fpga compile problem and kernels (#1989) · 0720653b
  由 TianXiaogang 提交于 9月 10, 2019
  
  0720653b
09 9月, 2019 3 次提交

J
add assign_value and hard_sigmoid, add fluid_type (#1983) · 92eeabeb
由 juncaipeng 提交于 9月 09, 2019
```
* add assign_value op, arm kernel and test, add fluid_type, test=develop

* add hard_sigmoid, test=develop
```
92eeabeb

Add concat and elementwise_add cuda kernel (#1979) · 6d1da405

由 Pei Yang 提交于 9月 09, 2019

* add nearest_interp_cuda kernel, test=develop

* add concat op and elementwise_add op

* remove eigen dependency from nearest_interp cuda kernel, test=develop

* free cuda pointers, test=develop

6d1da405

Z
add calib cuda kernel. (#1977) · da328594
由 Zhen Wang 提交于 9月 09, 2019
```
* add calib cuda kernel.

* add unit test for calib cuda kernel. test=develop
```
da328594

07 9月, 2019 1 次提交

add lite x86 ops for ASR test=develop (#1981) · 7014a76b

由 lijianshe02 提交于 9月 07, 2019

* add lite x86 ops for ASR test=develop

* add lite x86 ops for ASR test=develop

* fix x86 ci run test problems test=develop

* fix mkl path for CI test=develop

7014a76b

06 9月, 2019 2 次提交

H
modify reshape2 OP test=dvelop (#1963) · febfd7d6
由 huzhiqiang 提交于 9月 06, 2019
```
modify reshape2 OP to add shape_tensor input
```
febfd7d6

add cudnn conv fp32, int8 support (#1974) · f3124b30

由 Zhaolong Xing 提交于 9月 06, 2019

* paddle lite cuda init
can run model with leaky_relu

* add the missing file.
test=develop

* add the load from memory interface.
test=develop

* refine this pr. fix comments
fix ci error
test=develop

* conv impl
fp32:
conv, conv+bais, conv+bias+relu, conv+bias+leaky_relu

int8:
conv, conv+bais+relu(int8 or fp32 output), conv+bias+leaky_relu(int8 or fp32 output)

can run conv+ bias+relu using cxx_api
test=develop

* move the lite/cuda/math to backends/cuda/math
test=develop

f3124b30

03 9月, 2019 2 次提交

rewrite multiclass_nms according to fluid, test=develop (#1945) · deaddf9d

由 juncaipeng 提交于 9月 03, 2019

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

* add box_clip op and fix tensor slice bug

* remove add four op twice

* rewrite the implement for box_coder and sequence_expand, add faster_rcnn_test, test=develop

* fix test bug of box_clip in x86 server, test=develop

* rewrite multiclass_nms according to fluid, test=develop

* fix param load bug in box_coder and multiclass_nms op, test=develop

* fix value transfor error in multiclass_nms, test=develop

deaddf9d

H

create backends directory and move hardware backends into it (#1954) · 31ee212a
由 huzhiqiang 提交于 9月 03, 2019

31ee212a

02 9月, 2019 2 次提交

Add ops and fix bugs for Faster RCNN (#1942) · 635b4958

由 juncaipeng 提交于 9月 02, 2019

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

* add box_clip op and fix tensor slice bug

* remove add four op twice

* rewrite the implement for box_coder and sequence_expand, add faster_rcnn_test, test=develop

* fix test bug of box_clip in x86 server, test=develop

635b4958

fix elementwise_op bug when the shape of input y is 1, test=develop (#1924) · 6d9b1558

由 juncaipeng 提交于 9月 02, 2019

* fix elementwise_op bug when the shape of input y is 1, test=develop

* fix elementwise ops bug when the shape of input y is 1, test=develop

6d9b1558

30 8月, 2019 2 次提交
- L
  Modify the slice op, stack op, reduce_mean op (#1921) · 6dbc2f29
  由 liu zhengxi 提交于 8月 30, 2019
```
* add stack op and add reduce_mean op and their unit tests

* modify stack op output name and modify the for loop in reduce_mean op

* add HasAttr for slice op
```
  6dbc2f29
- P
  add nearest_interp_cuda kernel, test=develop (#1920) · 029971b4
  由 Pei Yang 提交于 8月 30, 2019
```
add nearest_interp cuda kernel for Paddle-Lite
```
  029971b4
29 8月, 2019 3 次提交

Add yolo_box_cuda multiclass_nms_host kernel. (#1908) · de43e479

由 Wilber 提交于 8月 29, 2019

* add yolo_box_compute cuda

* move multiclass_nms(arm) to host

* add lod in scale op

* add yolo_box_cuda cmake config

* modify shuffle_channel_fuse and transpose_softmax_transpose_fuse to support run ssd model. test=develop

* reshape and transpose op don't have xshape output.

* modify yolo_box_compute_cuda, use tensor to manage cuda memory test=develop

* add yolo_box use kernel test=develop

de43e479

L

add stack op and add reduce_mean op and their unit tests (#1888) · 20001636
由 liu zhengxi 提交于 8月 29, 2019

20001636

ad ops for faster rcnn, including affine_channel, anchor_generator,... · 53b05ce8

由 juncaipeng 提交于 8月 29, 2019

ad ops for faster rcnn, including affine_channel, anchor_generator, generate_proposals and roi_align (#1895)

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

53b05ce8

28 8月, 2019 1 次提交
- J
  Modify cast op and remove warning in argmax_test (#1894) · 0cfbd266
  由 juncaipeng 提交于 8月 28, 2019
```
* modify cast op, test=develop

* modify cast op and remove warning in argmax_test, test=develop
```
  0cfbd266