提交 · 95c1816ec0954fdfdce982eea85e65a30d8c3cb3 · Crayon鑫 / Paddle

23 7月, 2019 6 次提交

[MKL-DNN] Extended LRN with reusing via Acquire API (#18675) · 95c1816e

由 Jacek Czaja 提交于 7月 23, 2019

test=develop

- compileation fix

- Yet another compilation fix

- Even yet another compilation fix

- Surprise! Again compilation fix

- lint fixes

test=develop

- Fix to workspace acquire of LRN

test=develop

- Fix to hash of BWD LRN

test=develop

- fix to lrn BWD PD acquire

test=develop

- Fixing LRN PD creation

test=develop

- cosmetic fix in comment

test=develop

- Fixes after review

test=develop

95c1816e

T
remove unused cmake file (#18744) · 0ae45f0b
由 Tao Luo 提交于 7月 23, 2019
```
test=develop
```
0ae45f0b

support patch data, add load_one_table, fix bug (#18509) · d18aabb4

由 jiaqi 提交于 7月 23, 2019

（1）support patch data （merge slots of instances of same line id, modify dense layer which
changes its size）
（2）add fleet load_one_table interface, support load from paddle model and load from pslib model
（3）fix push sparse bug which cause push sparse cost more time（about 10% in my testcase）
（4）when some slots are not in one of your network (join/update, etc.)，data feed、collect label info、push/pull sparse will skip these slots， instead of throw error.
（5）add more debug info in TrainFilesWithProfiler

d18aabb4

C
Make fuse_optimizer_op_pass also work when the model contains sparse gradients. (#18664) · fd3aad6c
由 chengduo 提交于 7月 23, 2019
```
* support sparse gradients
test=develop
```
fd3aad6c

Cudnn convolution reconstruction (#18284) · 6b78e00d

由 wangchaochaohu 提交于 7月 23, 2019

* rewrite the conv_op using cudnn_conv_helper

* add workspace limit for v7 test=develop

* fix test=develop

* add half float test=develop

* fix test=develop

* fix test=develop

* revise code style test=develop

* fix test=develop

6b78e00d

supports distributed classification (#18690) · 157211c4

由 Yi Liu 提交于 7月 23, 2019

* supports distributed classification training
* update API.spec
* fix evenly division in python3
* change "index_range" to "index_num" in shard_index operator
test=document_preview
test=develop

157211c4

22 7月, 2019 11 次提交
- Q
  
  Fix CPU implementation of roi_align_op backward (#18728) · 3429e65a
  由 qingqing01 提交于 7月 22, 2019
  
  3429e65a
- G
  add parameter server launch (#18687) · 70b03760
  由 guru4elephant 提交于 7月 22, 2019
```
add parameter server launch so that a user can easily launch parameter server
```
  70b03760
- Z
  
  add more traceback to py_reader error msg, test=develop (#18722) · d07ad4c6
  由 Zeng Jinle 提交于 7月 22, 2019
  
  d07ad4c6
- H
  Fix random test_recurrent_op failure (#18718) · a3028bb7
  由 Huihuang Zheng 提交于 7月 22, 2019
```
The change includes 3 things:

1. Set CPU_NUM to 1 in the tests because the ParallelExecutor will print warning that CPU_NUM is not set and use default 1.

2. Old tests compare two RNNs, hand written simple RNN and same RNN built by Paddle, but initialized RNN weights in numpy random and Paddle random separately. Fixed it by setting weights and bias values.

3. Also set numpy random seed in the tests. Now the two RNNs diff can be smaller (rtol from 0.1, 0.2 to. 0.01) in the tests.

test=develop
```
  a3028bb7
- T
  Revert "Add LeakyRelu MKLDNN support (#18656)" (#18723) · bd22453f
  由 Tao Luo 提交于 7月 22, 2019
```
test=develop
```
  bd22453f
- T
  
  Change api approval people name (#18699) · 58469186
  由 tianshuo78520a 提交于 7月 22, 2019
  
  58469186
- W
  Make infer shape of pad2d support for input with negative dims in compile time. (#18695) · 189b08dc
  由 whs 提交于 7月 22, 2019
```
test=develop
```
  189b08dc
- T
  remove unused gzstream.cmake (#18705) · c457a69d
  由 Tao Luo 提交于 7月 22, 2019
```
test=develop
```
  c457a69d
- T
  do some odd jobs (#18641) · d8458483
  由 tangwei12 提交于 7月 22, 2019
```
do some odd jobs, test=develop
```
  d8458483
- B
  
  add license, test=develop (#18709) · 7e3963f2
  由 Bai Yifan 提交于 7月 22, 2019
  
  7e3963f2
- G
  split different comm method for mnist distributed training (#18715) · ebf9797e
  由 guru4elephant 提交于 7月 22, 2019
```
* split different comm method for mnist distributed training
```
  ebf9797e
20 7月, 2019 2 次提交
- C
  test=develop (#18701) · ccf06a48
  由 cjt222 提交于 7月 20, 2019
```
add license
```
  ccf06a48
- W
  fix clip_by_norm doc (#18688) · 185b3ace
  由 wangguanzhong 提交于 7月 20, 2019
```
* fix clip_by_norm doc, test=develop
```
  185b3ace
19 7月, 2019 5 次提交
- H
  Support memory eager deletion on recurrent OP (#17710) · 89bc3fd8
  由 Huihuang Zheng 提交于 7月 19, 2019
```
Test PaddingRNN on V100 GPU device.

Test configuration: large model, padding mode (which is the mode using recurrentOp), one GPU.
                   
GPU memory (MiB):   6414 (this PR)     vs   6837 (without this PR)
Speed (steps/s):         10.28 (this PR)    vs    9.89 (without this PR)
 
```
  89bc3fd8
- J
  MKL-DNN upgrade to 0.20 (#18370) · 0d8e6c9b
  由 Jacek Czaja 提交于 7月 19, 2019
```
test=develop
```
  0d8e6c9b
- A
  Add LeakyRelu MKLDNN support (#18656) · d6b6a337
  由 Adam 提交于 7月 19, 2019
```
test=develop
```
  d6b6a337
- T
  add check of executor (#17986) · 0b9acb49
  由 tangwei12 提交于 7月 19, 2019
```
* add check of executor, test=develop
```
  0b9acb49
- G
  
  Change to use brpc rdma branch instead of personal branch. (#18683) · ec1000cc
  由 gongweibao 提交于 7月 19, 2019
  
  ec1000cc
18 7月, 2019 6 次提交

Optimize the content of error reporting information, print error code and... · 772e0956

由 zhouwei25 提交于 7月 18, 2019

Optimize the content of error reporting information, print error code and official document web sites (#18671)

optimize the error reporting information of cuda related API
index on develop: 130ac177 Merge branch 'develop' of https://github.com/PaddlePaddle/Paddle into develop

772e0956

Feature/auto_growth_allocator (#18561) · ae58afc5

由 Zeng Jinle 提交于 7月 18, 2019

* feature/auto_growth_allocator, test=develop

* add unittest of AlignedAllocator, test=develop

* try to turn on auto_growth to test on CI, test=develop

* fix segmentation fault in mixed_vector.h, test=develop

* add unittests, test=develop

ae58afc5

H
hash_op support int64 hash_size (#18674) · bb2f5d24
由 hutuxian 提交于 7月 18, 2019
```
* hash_op support int64 hash_size
* add corresponding UT
```
bb2f5d24
X

update readme to 1.5.1 (#18670) · a5d4c2fa
由 xsrobin 提交于 7月 18, 2019

a5d4c2fa
G
remove ctr reader, all functions are satisfied in dataset (#18672) · 5ed713d5
由 guru4elephant 提交于 7月 18, 2019
```
* remove ctr reader, all functions are satisfied in dataset
```
5ed713d5

Downgrade gcc to 4.8 (#18614) · 898237c1

由 Jiabin Yang 提交于 7月 18, 2019

* test=develop, fix docker with paddle nccl problem

* test=develop, downgrade gcc to 4.8 for latest-dev

* test=develop, downgrade gcc to 4.8 for latest-dev

* test=develop, modify cmake to renew all third_party

* test=develop, invoke ci

* test=develop, invoke ci

* test=develop, complie python with wide-unicode

* test=deveop, refine env settings

* test=deveop, refine env settings

898237c1

17 7月, 2019 5 次提交

G
remove async executor and add data_feed.proto to the deps of train demo (#18659) · d714bf03
由 guru4elephant 提交于 7月 17, 2019
```
* remove async executor and add data_feed.proto to the deps of train demo
```
d714bf03

Add cuda implementation for `prelu` backward pass (#18633) · ce1ec332

由 Yang Zhang 提交于 7月 17, 2019

* Add GPU implementation for `prelu` backward pass

test=develop

* Fix logic error in `prelu` GPU backward and simplify a bit

test=develop

* Fix `prelu` backward CUDA implementation

test=develop

CPU version was not used actually, so test passed

ce1ec332

石

Fix Bitmain Predictor::Clone() (#18599) · 25d80791

由石晓伟提交于 7月 17, 2019

* update anakin-engine interfaces for content-dnn

test=develop

* support only-gpu mode of Anakin

modify eltwise parse

test=develop

* modification for thread-safe

test=develop

* Integrated template instance

test=develop

* increase template parameters

test=develop

* support MLU predictor

test=develop

* update anakin cmake files

test=develop

* update TargetWrapper::set_device

* update the initialization of anakin subgraph

test=develop

* use the default constructor of base class

test=develop

* load model from buffer with length

test=develop

* modify the access level of class

test=develop

* support anakin for bitmain arch

test=develop

* remove files

* checkout cmakelists

test=develop

* modify interfaces

test=develop

* add cmake dependments

test=develop

* enforce the outputs of net

test=develop

25d80791

Y

[CPU] Fix the compiling issue with AVX512F macro. (#18634) · 97549a4f
由 Yihua Xu 提交于 7月 17, 2019

97549a4f
B

[NGraph] handle dim element 0 of ngraph op (#18568) · 256ba7cb
由 baojun 提交于 7月 16, 2019

256ba7cb

16 7月, 2019 4 次提交

C
fix PE fetch bug (#18644) · a6d468a2
由 chengduo 提交于 7月 16, 2019
```
test=develop
```
a6d468a2
L

print out error code of cudaGetDeviceProperties if failed (#18643) · 75953096
由 liuwei1031 提交于 7月 16, 2019

75953096

[MKL-DNN] Reimplemented pool2d mkl-dnn to use Acquire API (#18585) · 71d883b8

由 Jacek Czaja 提交于 7月 16, 2019

* - Added partial draft of pooling acquire

- Workspace support

- compilation fix

- Added draft of pooling backward reimplementation

- Segfault fix

- reverted 'any' for diff_dst crewation in pooling

- Lint fixes

test=develop

- lint fixes

test=develop

- Further lint fixes

test=develop

* - Fixes after review

test=develop

* - Lint fixes

test=develop

* - Even more lint fixes

test=develop

71d883b8

C
fix bug of scatter op (#18640) · f4ec7d54
由 chengduo 提交于 7月 16, 2019
```
test=develop
```
f4ec7d54

15 7月, 2019 1 次提交
- T
  
  change pip install whl;test=develop (#18635) · 112cf850
  由 tianshuo78520a 提交于 7月 15, 2019
  
  112cf850

Crayon鑫 / Paddle 与 Fork 源项目一致

Crayon鑫 / Paddle
与 Fork 源项目一致