提交 · 04f4775b346be540f3b446074112d9bdcef903c7 · PaddlePaddle / Paddle-Lite

06 9月, 2019 4 次提交

Z
add interpolate fuse pass (#1980) · c49958a2
由 zhupengyang 提交于 9月 06, 2019
```
test=develop
```
c49958a2
H
modify reshape2 OP test=dvelop (#1963) · febfd7d6
由 huzhiqiang 提交于 9月 06, 2019
```
modify reshape2 OP to add shape_tensor input
```
febfd7d6
W
modify nearest_interpolate when attr align_corners=false (bug_fix) (#1969) · 4a3a45b6
由 Wilber 提交于 9月 06, 2019
```
* modify slice op and add slice test

* modify nearest_polate when align_corners=false (bugfix)
```
4a3a45b6

add cudnn conv fp32, int8 support (#1974) · f3124b30

由 Zhaolong Xing 提交于 9月 06, 2019

* paddle lite cuda init
can run model with leaky_relu

* add the missing file.
test=develop

* add the load from memory interface.
test=develop

* refine this pr. fix comments
fix ci error
test=develop

* conv impl
fp32:
conv, conv+bais, conv+bias+relu, conv+bias+leaky_relu

int8:
conv, conv+bais+relu(int8 or fp32 output), conv+bias+leaky_relu(int8 or fp32 output)

can run conv+ bias+relu using cxx_api
test=develop

* move the lite/cuda/math to backends/cuda/math
test=develop

f3124b30

04 9月, 2019 1 次提交
- W
  modify slice op and add slice test (#1944) · cfc7af76
  由 Wilber 提交于 9月 04, 2019
```
* modify slice op and add slice test

* modify slice op bug
```
  cfc7af76
03 9月, 2019 6 次提交

refine the npu graph and subgraph (#1959) · d60b8d61

由 tensor-tang 提交于 9月 03, 2019

* fix attr and refine subgraph pass test=develop

* refine the npu pass functions

* fix test

test=develop

d60b8d61

Z
enhance interpolate op when there is no "scale" (#1957) · 3191ec5e
由 zhupengyang 提交于 9月 03, 2019
```
test=develop
```
3191ec5e
H

move npu into backends(directory) and move python/ into tools/python (#1958) · c5e65402
由 huzhiqiang 提交于 9月 03, 2019

c5e65402

rewrite multiclass_nms according to fluid, test=develop (#1945) · deaddf9d

由 juncaipeng 提交于 9月 03, 2019

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

* add box_clip op and fix tensor slice bug

* remove add four op twice

* rewrite the implement for box_coder and sequence_expand, add faster_rcnn_test, test=develop

* fix test bug of box_clip in x86 server, test=develop

* rewrite multiclass_nms according to fluid, test=develop

* fix param load bug in box_coder and multiclass_nms op, test=develop

* fix value transfor error in multiclass_nms, test=develop

deaddf9d

H

create backends directory and move hardware backends into it (#1954) · 31ee212a
由 huzhiqiang 提交于 9月 03, 2019

31ee212a
T
refine the attr assert (#1947) · 60bbc691
由 tensor-tang 提交于 9月 03, 2019
```
test=develop
```
60bbc691

02 9月, 2019 3 次提交

Y
[LITE][BENCHMARK] enhance android arm cpu benchmark (#1939) · ae6c1704
由 Yuan Shuai 提交于 9月 02, 2019
```
* enhance benchmark
* update code format. test=develop
```
ae6c1704

Add ops and fix bugs for Faster RCNN (#1942) · 635b4958

由 juncaipeng 提交于 9月 02, 2019

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

* add box_clip op and fix tensor slice bug

* remove add four op twice

* rewrite the implement for box_coder and sequence_expand, add faster_rcnn_test, test=develop

* fix test bug of box_clip in x86 server, test=develop

635b4958

fix elementwise_op bug when the shape of input y is 1, test=develop (#1924) · 6d9b1558

由 juncaipeng 提交于 9月 02, 2019

* fix elementwise_op bug when the shape of input y is 1, test=develop

* fix elementwise ops bug when the shape of input y is 1, test=develop

6d9b1558

01 9月, 2019 2 次提交

[ARM][CPU] Fix time counter of arm cpu profiler (#1925) · e3fb95ae

由 Yuan Shuai 提交于 9月 01, 2019

* Fix timer of arm cpu profiler. test=develop

* Fix un-added op in cmake.test=develop

* fix cmake error

* fix cmake error, test=develop

* Fix pass sequence. test=develop

* replace option with lite_option. test=develop

* disable profile mode by default. test=develop

* Fix error option name. test=develop

e3fb95ae

H

add the method of loading model from naive buffer for LightPredictor (#1918) · 13715c57
由 huzhiqiang 提交于 9月 01, 2019

13715c57

31 8月, 2019 1 次提交
- G
  
  lite can build armlinux and remove build_armlinux.sh test=develop (#1932) · f1252590
  由 guofei 提交于 8月 31, 2019
  
  f1252590
30 8月, 2019 6 次提交

Modify the slice op, stack op, reduce_mean op (#1921) · 6dbc2f29

由 liu zhengxi 提交于 8月 30, 2019

* add stack op and add reduce_mean op and their unit tests

* modify stack op output name and modify the for loop in reduce_mean op

* add HasAttr for slice op

6dbc2f29

P
add nearest_interp_cuda kernel, test=develop (#1920) · 029971b4
由 Pei Yang 提交于 8月 30, 2019
```
add nearest_interp cuda kernel for Paddle-Lite
```
029971b4

[NPU] add NPU supporting for Java API (#1915) · dbabf5c4

由 hong19860320 提交于 8月 30, 2019

* [NPU] add NPU supporting for Java API
test=develop

* [NPU] refine build script for NPU compiling
test=develop

* [NPU] fix compiling script for NPU
test=develop

dbabf5c4

Z
add npu pad2d op converter (#1896) · b1a8c2a2
由 zhupengyang 提交于 8月 30, 2019
```
test=develop
```
b1a8c2a2

support ios tiny publish (#1910) · d8ccafcd

由 Xiaoyang LI 提交于 8月 30, 2019

* fix ios build script, test=develop

* add ios tiny publish target, test=develop

* fix ios build scrip, test=develop

* merge build_ios.sh to build.sh, test=develop

* ios support BUILD_EXTRA, test=develop

d8ccafcd

add precision and persistable attrs for the tensor. (#1899) · e2e07fa4

由 Zhen Wang 提交于 8月 30, 2019

* Add precision and persistable attrs for the tensor. And fix cxx light and full api demo.

* update precision2string methods. test=develop

* move the save logic to the front of the run in mobilenetv1_full_api.cc, test=develop.

* add comments for UpdateVarsOfProgram. test=develop

e2e07fa4

29 8月, 2019 10 次提交

T

[NPU] fix npu compile of publish_inference (#1911) · 443b2e68
由 tensor-tang 提交于 8月 29, 2019

443b2e68
T

add conv2d transpose fuse (#1909) · 57ee8714
由 tensor-tang 提交于 8月 29, 2019

57ee8714

Add yolo_box_cuda multiclass_nms_host kernel. (#1908) · de43e479

由 Wilber 提交于 8月 29, 2019

* add yolo_box_compute cuda

* move multiclass_nms(arm) to host

* add lod in scale op

* add yolo_box_cuda cmake config

* modify shuffle_channel_fuse and transpose_softmax_transpose_fuse to support run ssd model. test=develop

* reshape and transpose op don't have xshape output.

* modify yolo_box_compute_cuda, use tensor to manage cuda memory test=develop

* add yolo_box use kernel test=develop

de43e479

S

[Java API][Comment] upate java api & delete some comments (#1912) · 79714d74
由 sangoly 提交于 8月 29, 2019

79714d74
L

add stack op and add reduce_mean op and their unit tests (#1888) · 20001636
由 liu zhengxi 提交于 8月 29, 2019

20001636

Add load from memory interface (#1903) · ecce1eff

由 Zhaolong Xing 提交于 8月 29, 2019

* paddle lite cuda init
can run model with leaky_relu

* add the missing file.
test=develop

* add the load from memory interface.
test=develop

* refine this pr. fix comments
fix ci error
test=develop

ecce1eff

S

[Java API] add setThreads & setPowerMode interface (#1907) · e91eef1c
由 sangoly 提交于 8月 29, 2019

e91eef1c
T
[NPU] enable npu program rollback (#1906) · 178a93b9
由 tensor-tang 提交于 8月 29, 2019
```
test=develop
```
178a93b9

ad ops for faster rcnn, including affine_channel, anchor_generator,... · 53b05ce8

由 juncaipeng 提交于 8月 29, 2019

ad ops for faster rcnn, including affine_channel, anchor_generator, generate_proposals and roi_align (#1895)

* add ops for faster rcnn

* disable test for generate_proposals and roi_align, test=develop

* remove .swp file

* remove log in tensor slice

* finish the unit test for roi_align, test=develop

53b05ce8

[NPU] refine npu subgraph and clean code (#1902) · 0c25428c

由 tensor-tang 提交于 8月 29, 2019

* add npu script and tester

* fix npu armv7 so and refine tests

test=develop

* update fix and refine log

test=develop

* refine npu generate api

* refine npu subgraph

* refine npu gen and clean code

* fix model laod

* refine node2rm in subgraph

* refine the build npu functions

test=develop

0c25428c

28 8月, 2019 7 次提交
- J
  Modify cast op and remove warning in argmax_test (#1894) · 0cfbd266
  由 juncaipeng 提交于 8月 28, 2019
```
* modify cast op, test=develop

* modify cast op and remove warning in argmax_test, test=develop
```
  0cfbd266
- Y
  
  change num proc; test=develop (#1889) · 788238b0
  由 Yan Chunwei 提交于 8月 28, 2019
  
  788238b0
- Z
  add nearest_interp op converter (#1879) · c3d15d76
  由 zhupengyang 提交于 8月 28, 2019
```
test=developt branch
```
  c3d15d76
- H
  
  add floor op,elementwise_div op and assign op test=develop (#1882) · 26450c49
  由 huzhiqiang 提交于 8月 28, 2019
  
  26450c49
- Z
  add transpose-softmax-transpose fuse pass (#1863) · 5e8b15f5
  由 zhupengyang 提交于 8月 28, 2019
```
* add transpose-softmax-transpose fuse pass

test=develop

* enable supported lite-npu ops

test=develop
```
  5e8b15f5
- H
  add x86 math:sequence_scale,sequence_padding,sequence2batch,sequence_pooling. test=develop (#1884) · 54101ef0
  由 huzhiqiang 提交于 8月 28, 2019
```
add x86 math:sequence_scale,sequence_padding,sequence2batch,sequence_pooling. test=develop (#1884)
```
  54101ef0
- H
  [NPU] fix conv2d npu bridge, supports bias from input map (#1839) · 1ee60474
  由 hong19860320 提交于 8月 28, 2019
```
* [NPU] fix conv2d npu bridge, supports bias from input map
test=develop

* [NPU] support more dimensions for the bias of conv2d NPU bridge
test=develop
```
  1ee60474