提交 · 61221ebc28a9cc6f953715b54838b810e06f8df9 · PaddlePaddle / PaddleDetection

25 5月, 2019 1 次提交

TRT: Support set dynamic range in int8 mode. (#17524) · 61221ebc

由 Zhaolong Xing 提交于 5月 25, 2019

* fluid int8 train and trt int8 predict align.
trt int8 predict init
op converter

* 2. align fluid int8 train and trt int8 inference.
enhance quant dequant fuse pass
enhance op converter, trt engine, trt engine op, trt subgraph pass.

* 3. add delete_quant_dequant_pass for trt

test=develop

* 4. add the missing file
test=develop

* 5. i modify the c++ interface, but forget to modify the pybind code
fix the IS_TRT_VERSION_GE bug, and fix elementwise op converter
test=develop

61221ebc

24 5月, 2019 2 次提交

[MKL-DNN] Add Fully Connected Op for inference only(#15226) · 0c39b97b

由 Michał Gallus 提交于 5月 24, 2019

* fuse mul and elementwise add to fc

* Reimplement the FC forward operator

* Fix FC MKLDNN integration by transposing weights

* Add FC MKLDNN Pass

test=develop

* FC MKLDNN Pass: change memcpy to std::copy

* Fix MKLDNN FC handling of mismatch input and weights dims

* Lower tolerance for MKL-DNN in resnet50 test

test=develop

* Adjust FC to support MKLDNN Op placement

test=develop

* Adjust Placement Op to set use_mkldnn attribute for graph

test=develop

* MKLDNN FC: fix weights format so that gemm version is called

test=develop

* FC MKLDNN: Remove tolerance decrease from tester_helper

* FC MKL-DNN: Refactor the code, change input reorder to weight reorder

* MKL-DNN FC: Introduce operator caching

test=develop

* FC MKL-DNN: Fix the tensor type in ExpectedKernelType

test=develop

* FC MKL-DNN: fix style changes

test=develop

* FC MKL-DNN: fallback to native on non-supported dim sizes

test=develop

* FC MKLDNN: fix CMake paths

test=develop

* FC MKLDNN: Refine placement pass graph mkldnn attribute

test=develop

* Fix Transpiler error for fuse_conv_eltwise

test=develop

* Fix missing STL includes in files

test=develop

* FC MKL-DNN: Enable new output size computation

Also, refine pass to comply with newest interface.
test=develop

* FC MKL-DNN: enable only when fc_mkldnn_pass is enabled

* FC MKL-DNN: Allow Weights to use oi or io format

* FC MKL-DNN: Adjust UT to work with correct dims

test=develop

* Enable MKL DEBUG for resnet50 analyzer

test=develop

* FC MKL-DNN: Improve Hashing function

test=develop

* FC MKL-DNN: Fix shape for fc weights in transpiler

* FC MKL-DNN: Update input pointer in re-used fc primitive

* Add log for not handling fc fuse for unsupported dims

test=develop

* FC MKL-DNN: Move transpose from pass to Op Kernel

test=develop

* FC MKL-DNN: Disable transpose in unit test

test=develop

* FC MKL-DNN: Remove fc_mkldnn_pass from default list

* Correct Flag for fake data analyzer tests

test=develop

* FC MKL-DNN: Add comment about fc mkldnn pass disablement

test=develop

* FC MKL-DNN: Disable fc in int8 tests

test=develop

0c39b97b

Conv concat relu quantization (#17466) · 5b2a3c4b

由 Sylwester Fraczek 提交于 5月 24, 2019

* add conv_concat_relu fuse

test=develop

* add test code

test=develop

* added missing include with unordered_map

test=develop

* review fixes for wojtuss

test=develop

* remove 'should (not) be fused' comment statements

one of them was invalid anyway

test=develop

5b2a3c4b

22 5月, 2019 1 次提交

Enable the convolution/relu6(bounded_relu) fusion for FP32 on Intel platform. (#17130) · 2281ebf0

由 guomingz 提交于 5月 22, 2019

* Relu6 is the bottleneck op for Mobilenet-v2. As the mkldnn supports the conv/relu6 fusion, we implement it fusion via cpass way. Due to the int8 enabling for this fusion will be supported in MKLDNN v0.20, so this PR is focused on the fp32 optimization.

Below table shows the benchmark(FPS) which measured on skx-8180(28 cores)
Batch size | with fusion | without fusion
-- | -- | --
1 | 214.7 | 53.4
50 | 1219.727 | 137.280

test=develop

* Fix the format issue

test=develop

* Add the missing nolint comments.

test=develop

* Fix the typos.

test=develop

* Register the conv_brelu_mkldnn_fuse_pass for the MKLDNN engine.

test=develop

* Adjust the indentation.

test=develop

* Add the test_conv_brelu_mkldnn_fuse_pass case.

test=develop

* Slightly update the code per Baidu comments.
Let the parameter definition embedded into the code.
That's will make the code easy to understand.

test=develop

2281ebf0

21 5月, 2019 1 次提交

fix security bugs : (#17464) · ba70cc49

由 liuwei1031 提交于 5月 21, 2019

http://newicafe.baidu.com:80/issue/PaddleSec-33/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-28/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-25/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-24/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-21/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-20/show?from=page

test=develop

ba70cc49

20 5月, 2019 2 次提交
- T
  remove unused expected_kernel_cache_pass (#17486) · 32da5e9c
  由 Tao Luo 提交于 5月 20, 2019
```
test=develop
```
  32da5e9c
- W
  fix the random compilation failure on windows test=develop (#17475) · ca3ba378
  由 wopeizl 提交于 5月 20, 2019
```
* fix the random compilation failure on windows 
```
  ca3ba378
16 5月, 2019 1 次提交

Add setting Scope function for the graph class (#17417) · 4a1b7fec

由 Zhen Wang 提交于 5月 16, 2019

* add set_not_owned function for graph

* add scope set. test=develop

* add scope_ptr enforce not null before setting.test=develop

4a1b7fec

09 5月, 2019 1 次提交

fix: (#17279) · 7a3bb061

由 Zhaolong Xing 提交于 5月 09, 2019

1. infernce multi card occupy
2. facebox model inference occupy too much
test=develop

7a3bb061

08 5月, 2019 1 次提交
- W
  improved unit test output (#17266) · 984aa905
  由 Wojciech Uss 提交于 5月 08, 2019
```
added printing data type to differentiate int8 and fp32 latency results

test=develop
```
  984aa905
07 5月, 2019 2 次提交

石

Cherry-pick benchmark related changes from release/1.4 (#17156) · a72dbe9a

由石晓伟提交于 5月 07, 2019

* cherry-pick commit from 88770542

* cherry-pick commit from 3f0b97df

* cherry-pick from 16691:Anakin subgraph support yolo_v3 and faster-rcnn

(cherry picked from commit 8643dbc2)

* Cherry-Pick from 16662 : Anakin subgraph cpu support

(cherry picked from commit 7ad182e1)

* Cherry-pick from 1662, 16797.. : add anakin int8 support

(cherry picked from commit e14ab180)

* Cherry-pick from 16813 : change singleton to graph RegistBlock
test=release/1.4

(cherry picked from commit 4b9fa423)

* Cherry Pick : 16837 Support ShuffleNet and MobileNet-v2

Support ShuffleNet and MobileNet-v2, test=release/1.4

(cherry picked from commit a6fb066f)

* Cherry-pick : anakin subgraph add opt config layout argument #16846
test=release/1.4

(cherry picked from commit 8121b3ec)

* 1. add shuffle_channel_detect

(cherry picked from commit 6efdea89)

* update shuffle_channel op convert, test=release/1.4

(cherry picked from commit e4726a06)

* Modify symbol export rules

test=develop

a72dbe9a

call SetNumThreads everytime to avoid missing omp thread setting (#17224) · 54636a19

由 Leo Zhao 提交于 5月 07, 2019

* call SetNumThreads everytime to avoid missing omp thread setting

resolve #17153
test=develop

* add paddle_num_threads into config for test_analyzer_pyramid_dnn

resolve #17153
test=develop

54636a19

05 5月, 2019 1 次提交
- W
  
  use two GPUs to run the exclusive test test=develop (#17187) · 83c4f772
  由 wopeizl 提交于 5月 05, 2019
  
  83c4f772
23 4月, 2019 1 次提交
- L
  fix runtime_context_cache bug when gpu model has an op runs only on cpu · 490e7462
  由 luotao1 提交于 4月 23, 2019
```
test=develop
```
  490e7462
22 4月, 2019 1 次提交

add parallel build script to ci … (#16901) · d9991dcc

由 wopeizl 提交于 4月 22, 2019

* add parallel build script to ci test=develop
* 1. classify the test case as single card/two cards/multiple cards type
   2. run test case according to the run type

d9991dcc

19 4月, 2019 1 次提交
- T
  disable runtime_context_cache pass by default · aa7b975b
  由 Tao Luo 提交于 4月 19, 2019
```
test=develop
```
  aa7b975b
15 4月, 2019 1 次提交
- L
  add SaveOptimModel interface in analysis_predictor.h and test it in a… (#16441) · de26df44
  由 lijianshe02 提交于 4月 15, 2019
```
* add SaveOptimModel interface in analysis_predictor.h and test it in analyzer_dam_tester and analyzer_resnet50_tester test=develop
```
  de26df44
12 4月, 2019 1 次提交
- S
  fix memory optim temporarily · f58c3ec1
  由 superjomn 提交于 4月 12, 2019
```
test=develop
```
  f58c3ec1
11 4月, 2019 1 次提交

Security issue (#16774) · 85363848

由 liuwei1031 提交于 4月 11, 2019

* disable memory_optimize and inpalce strategy by default, test=develop

* fix security issue
http://newicafe.baidu.com:80/issue/PaddleSec-3/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-8/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-12/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-32/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-35/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-37/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-40/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-43/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-44/show?from=page
http://newicafe.baidu.com:80/issue/PaddleSec-45/show?from=page

test=develop

* revert piece.cc, test=develop

* adjust api.cc,test=develop

85363848

09 4月, 2019 1 次提交
- T
  disable seqpool concat pass by default saving CI time · d6c1b5a7
  由 tensor-tang 提交于 4月 09, 2019
```
test=develop
```
  d6c1b5a7
03 4月, 2019 3 次提交
- L
  test_analyzer_int8 tests use default pass order · bd636a9e
  由 luotao1 提交于 4月 03, 2019
```
test=develop
```
  bd636a9e
- L
  
  merge confict, test=develop · b236091e
  由 lujun 提交于 4月 03, 2019
  
  b236091e
- Y
  
  fix identity temporarily (#15942) · 044ae249
  由 Yan Chunwei 提交于 4月 03, 2019
  
  044ae249
02 4月, 2019 2 次提交
- W
  
  fix repeating passes (#16606) · ec2750b3
  由 Wojciech Uss 提交于 4月 02, 2019
  
  ec2750b3
- W
  
  fix dataset reading and add support for full dataset (#16559) · 9b6a0296
  由 Wojciech Uss 提交于 4月 02, 2019
  
  9b6a0296
29 3月, 2019 1 次提交
- S
  
  resolve conflicts with the develop branch test=develop · bddb2cd3
  由 Shixiaowei02 提交于 3月 28, 2019
  
  bddb2cd3
28 3月, 2019 3 次提交

Anakin ssd support · d065b5bf

由 nhzlx 提交于 3月 28, 2019

refine trt first run
add quant dequant fuse pass
omit simplify_anakin_priorbox_detection template
omit transpose_flatten_concat_fuse template
test=develop

d065b5bf

Refine default MKL-DNN Pass order (#16490) · 2d8b7b3a

由 Michał Gallus 提交于 3月 27, 2019

* Refine default MKL-DNN Pass order

test=develop

* Add comment to default MKL-DNN Pass list

test=develop

2d8b7b3a

C-API quantization core 2 (#16396) · 09dfc7a2

由 Wojciech Uss 提交于 3月 27, 2019

* C-API quantization core

test=develop
Co-authored-by: NSylwester Fraczek <sylwester.fraczek@intel.com>

* Decouple Quantizer from AnalysisPredictor

test=develop

* fixes after review

test=develop

* renamed mkldnn quantize stuff

test=develop

* remove ifdef from header file

test=develop

09dfc7a2

26 3月, 2019 1 次提交
- N
  fix comments · 45b3766f
  由 nhzlx 提交于 3月 26, 2019
```
test=develop
```
  45b3766f
25 3月, 2019 1 次提交
- L
  fix cdn issue, test=develop (#16423) · de3b70a1
  由 liuwei1031 提交于 3月 25, 2019
```
* fix cdn issue, test=develop

* fix cdn issue, test=develop
```
  de3b70a1
22 3月, 2019 1 次提交
- N
  1. Add ANAKIN_ROOT compile option · f3a2e4b3
  由 nhzlx 提交于 3月 22, 2019
```
2. refine trt code
test=develop
```
  f3a2e4b3
21 3月, 2019 2 次提交
- L
  add expected_kernel_cache_pass · 056599a7
  由 luotao1 提交于 3月 21, 2019
```
test=develop
```
  056599a7
- W
  Add enabling quantization (#16326) · cbe2dbf0
  由 Wojciech Uss 提交于 3月 21, 2019
```
* Add enabling quantization

test=develop

* remove unused (here) function
```
  cbe2dbf0
20 3月, 2019 6 次提交
- N
  cherry-pick from feature/anakin-engine: add data type for zero copy #16313 · 4f4daa4b
  由 nhzlx 提交于 3月 20, 2019
```
1. refine anakin engine
2. add data type for zero copy

align dev branch and PaddlePaddle:feature/anakin-engine brach
the cudnn workspace modify was not included for now, because we use a hard code way
in feature/anakin-engine branch. There should be a better way to implement it,
and subsequent submissions will be made.

test=develop
```
  4f4daa4b
- N
  
  git cherry-pick from feature/anakin-engine: update anakin subgraph #16278 · 07dcf285
  由 nhzlx 提交于 3月 20, 2019
  
  07dcf285
- N
  
  cherry-pick from feature/anakin-engine: refine paddle-anakin to new interface. #16276 · c407dfa3
  由 nhzlx 提交于 3月 20, 2019
  
  c407dfa3
- N
  
  cherry-pick from feature/anakin-engine: deal the changing shape when using anakin #16189 · a25331bc
  由 nhzlx 提交于 3月 20, 2019
  
  a25331bc
- N
  
  cherry-pick from feature/anakin-engine: add batch interface for pd-anakin #16178 · c79f06d3
  由 nhzlx 提交于 3月 20, 2019
  
  c79f06d3
- N
  cherry-pick from feature/anakin-engine: refine anakin subgraph. #16157 · 69d37f81
  由 nhzlx 提交于 3月 20, 2019
```
support change input size
```
  69d37f81

PaddlePaddle / PaddleDetection 大约 1 年 前同步成功

PaddlePaddle / PaddleDetection
大约 1 年前同步成功