提交 · 220eef602e0e0af0d5566231397766cff5d080db · Crayon鑫 / Paddle

24 7月, 2019 3 次提交

Extend Matmul to support matrix multiplication with multiple heads (#18570) · 220eef60

由 Bob Zhu 提交于 7月 24, 2019

* extend matmul op to support multiple head multiplication

With the support of multiple head, the multiplication of two big matrixes is
split into multiplication of several (head_number) small matrixes. e.g. if
Mat A is [3, 24] and Mat B is [24, 4], when multiple A and B with head_number
as 4, Mat A will be split as 4 matrix of [3, 6] and Mat B will be 4 matrix of
[6, 4]. The result of final matrix will be 4 matrix of [3, 4], i.e. [3, 16].

220eef60

Add python API for appending LoD level (#18702) · 075e1cf7

由 whs 提交于 7月 24, 2019

* Make lod reset op support for append lod level.

* Fix API.spec
test=develop

* Fix unitest.
test=develop

* Add python api for lod append.
test=develop

* Fix API.spec
test=develop

* Fix format of doc.
test=develop

* Fix unitest.
test=develop

* Fix doc.
test=develop

075e1cf7

C
Enhance backward process (#18700) · 8259f141
由 chengduo 提交于 7月 24, 2019
```
* prun backward ops
test=develop
```
8259f141

23 7月, 2019 2 次提交
- C
  Make fuse_optimizer_op_pass also work when the model contains sparse gradients. (#18664) · fd3aad6c
  由 chengduo 提交于 7月 23, 2019
```
* support sparse gradients
test=develop
```
  fd3aad6c
- Y
  supports distributed classification (#18690) · 157211c4
  由 Yi Liu 提交于 7月 23, 2019
```
* supports distributed classification training
* update API.spec
* fix evenly division in python3
* change "index_range" to "index_num" in shard_index operator
test=document_preview
test=develop
```
  157211c4
22 7月, 2019 3 次提交

Fix random test_recurrent_op failure (#18718) · a3028bb7

由 Huihuang Zheng 提交于 7月 22, 2019

The change includes 3 things:

1. Set CPU_NUM to 1 in the tests because the ParallelExecutor will print warning that CPU_NUM is not set and use default 1.

2. Old tests compare two RNNs, hand written simple RNN and same RNN built by Paddle, but initialized RNN weights in numpy random and Paddle random separately. Fixed it by setting weights and bias values.

3. Also set numpy random seed in the tests. Now the two RNNs diff can be smaller (rtol from 0.1, 0.2 to. 0.01) in the tests.

test=develop

a3028bb7

T
Revert "Add LeakyRelu MKLDNN support (#18656)" (#18723) · bd22453f
由 Tao Luo 提交于 7月 22, 2019
```
test=develop
```
bd22453f
G
split different comm method for mnist distributed training (#18715) · ebf9797e
由 guru4elephant 提交于 7月 22, 2019
```
* split different comm method for mnist distributed training
```
ebf9797e

19 7月, 2019 3 次提交

Support memory eager deletion on recurrent OP (#17710) · 89bc3fd8

由 Huihuang Zheng 提交于 7月 19, 2019

Test PaddingRNN on V100 GPU device.

Test configuration: large model, padding mode (which is the mode using recurrentOp), one GPU.

GPU memory (MiB): 6414 (this PR) vs 6837 (without this PR)
Speed (steps/s): 10.28 (this PR) vs 9.89 (without this PR)

89bc3fd8

A
Add LeakyRelu MKLDNN support (#18656) · d6b6a337
由 Adam 提交于 7月 19, 2019
```
test=develop
```
d6b6a337
T
add check of executor (#17986) · 0b9acb49
由 tangwei12 提交于 7月 19, 2019
```
* add check of executor, test=develop
```
0b9acb49

18 7月, 2019 2 次提交

Feature/auto_growth_allocator (#18561) · ae58afc5

由 Zeng Jinle 提交于 7月 18, 2019

* feature/auto_growth_allocator, test=develop

* add unittest of AlignedAllocator, test=develop

* try to turn on auto_growth to test on CI, test=develop

* fix segmentation fault in mixed_vector.h, test=develop

* add unittests, test=develop

ae58afc5

H
hash_op support int64 hash_size (#18674) · bb2f5d24
由 hutuxian 提交于 7月 18, 2019
```
* hash_op support int64 hash_size
* add corresponding UT
```
bb2f5d24

15 7月, 2019 2 次提交
- G
  make auc op compatible with 1 dim (#18551) · ab57d389
  由 guru4elephant 提交于 7月 15, 2019
```
* make auc op compatible with 1 dim
```
  ab57d389
- G
  increase timeout again (#18628) · b71b4543
  由 guru4elephant 提交于 7月 15, 2019
```
test=develop
```
  b71b4543
12 7月, 2019 2 次提交
- 1
  fix #17430: int64类型的attr训练非预期 (#18264) · b414645a
  由 123malin 提交于 7月 12, 2019
```
* fix int64_t

* update fill constant op unittest

* add empty line
```
  b414645a
- K
  1）change to parallel mode on python coverage run (#18594) · 9ad57f2d
  由 kh2se2013 提交于 7月 12, 2019
```
2）add pip install coverage in Dockerfile.tmp
test=develop
```
  9ad57f2d
11 7月, 2019 2 次提交

G

Polish backwards optimizer dependency codes and use more default values. (#18255) · c0a82748
由 gongweibao 提交于 7月 11, 2019

c0a82748

Feature/buffer_shared_inplace (#17911) · d3003a16

由 Zeng Jinle 提交于 7月 11, 2019

* feature/buffer_shared_inplace, test=develop

* refine code, test=develop

* fix elementwise_add op cpu inplace and sum inplace bug, test=develop

* add unittest and debug log, test=develop

* fix parallel_executor scope bug, polish code, test=develop

* fix sum op, activation op, single_in_place_inference bug, test=develop

* remove kLocalExecScopeName, test=develop

* fix unittest,test=develop

* fix out_var first version bug, test=develop

* follow comments,test=develop

d3003a16

10 7月, 2019 1 次提交
- L
  update dygraph api doc for web (#18550) · b6d5c74f
  由 lujun 提交于 7月 10, 2019
```
remove dygraph.enable from __all__
hidden dygraph. profiler
add doc to dygraph. no_grad
```
  b6d5c74f
09 7月, 2019 2 次提交
- P
  
  Add mkldnn int8 mul-op kernel (#17834) · 0caa08ea
  由 Physher 提交于 7月 09, 2019
  
  0caa08ea
- L
  Fix roi_perspective_transform_op bug (#18522) · 24d1c44a
  由 LielinJiang 提交于 7月 09, 2019
```
* fix transform matrix bug, test=develop

* modify API.spec
```
  24d1c44a
05 7月, 2019 3 次提交

Fix topk cannot handle 1D vector bug (#18466) · 832d8191

由 zhaoyuchen2018 提交于 7月 05, 2019

* Fix topk cannot handle 1D vector bug

Add path to handle 1D vector

test=develop
Signed-off-by: Nzhaoyuchen <zhaoyuchen01@baidu.com>

* refine code

test=develop
Signed-off-by: Nzhaoyuchen <zhaoyuchen01@baidu.com>

832d8191

Hide no support (#18515) · 7586cdd5

由 Jiabin Yang 提交于 7月 05, 2019

* test=develop, fix docker with paddle nccl problem

* test=develop, hide no_support api and add ut for it

7586cdd5

Add distributions of normal and uniform (#18023) · 43e17c79

由 LielinJiang 提交于 7月 05, 2019

* add_distributions_of_normal_and_uniform

* paddle/fluid/API.spec

* modify API.spec

* modified paddle/fluid/API.spec, test=develop

* modify paddle/fluid/API.spec, test=develop

* modify paddle/fluid/API.spec, test=develop

* fix some comment, test=develop

* modify API.spec, test=develop

* add comment for init function, modify hard code, test=develop

* modify API.spec, test=develop

* modify API.spec, test=develop

* make unit test function shorter, test=develop

* modify paddle/fluid/API.spec

43e17c79

04 7月, 2019 2 次提交
- Q
  Enhance linear_lr_warmup (#18463) · 602cb6a5
  由 qingqing01 提交于 7月 04, 2019
```
* make it support float/int learning as input.
```
  602cb6a5
- C
  
  Make fuse_all_reduce_op_pass support mix_precision (#17652) · 74538573
  由 chengduo 提交于 7月 04, 2019
  
  74538573
03 7月, 2019 7 次提交
- Z
  
  support Tensor input for edit_distance op (#18162) · 7c6f2350
  由 zhoukunsheng 提交于 7月 03, 2019
  
  7c6f2350
- Z
  support Tensor input for chunk_eval op (#18226) · 26318544
  由 zhoukunsheng 提交于 7月 03, 2019
```
* test=develop
support Tensor input for chunk_eval op

* test=develop
fix testcase for chunk_eval op

* test=develop
fix typos in nn.py
```
  26318544
- Z
  
  add unique kernel and op (#17557) · 206c44e2
  由 zhoukunsheng 提交于 7月 03, 2019
  
  206c44e2
- Z
  
  upgrade hash op to support Tensor and LoDTensor input (#17998) · 71af72b1
  由 zhoukunsheng 提交于 7月 03, 2019
  
  71af72b1
- Z
  
  add ones_like op (#17388) · d3b3443d
  由 zhoukunsheng 提交于 7月 03, 2019
  
  d3b3443d
- Z
  
  add size op (#17412) · 67b48d7f
  由 zhoukunsheng 提交于 7月 03, 2019
  
  67b48d7f
- H
  Refactor for Pipeline Thread Check (#18459) · 6e0df310
  由 hutuxian 提交于 7月 03, 2019
```
move the thread-check code from train_from_dataset to a single function
add UT for the thread check function
```
  6e0df310
02 7月, 2019 2 次提交

supports collective training with programs (#18392) · a873fa84

由 Yi Liu 提交于 7月 02, 2019

1. Since allreduce op has 4 reduce types, We split these four reduce types into four ops
2. We also refined the collective op code, e.g. we separated the collective op kernel into CPUKernel and CUDAKernel, and remove the device specified DeviceContext parameter in template as we already knew the target DeviceContext
3. We remove the newly added Collective op role to reduce the complexity of program and graph analysis

a873fa84

C
Add find_no_grad_vars in backward.py (#17942) · e0d8c6ac
由 chengduo 提交于 7月 02, 2019
```
* add not_been_used_vars to no_grad_set
test=develop
```
e0d8c6ac

01 7月, 2019 1 次提交

Make roi_perspective_transform op return mask and transform matrix (#18371) · 449c7a9f

由 LielinJiang 提交于 7月 01, 2019

* modify roi_perspective_transform_op to output mask and transform matrix

* modify comment

* modify comment

* modify API.spec

* update API.spec

* remove no use header, test=develop

* resolve conflict

449c7a9f

27 6月, 2019 2 次提交

add WITH_COVERAGE option, default OFF (#17872) · 27fb9cad

由 kh2se2013 提交于 6月 27, 2019

* add WITH_COVERAGE option, default OFF

test=develop

* add coverage for python sdk

test=develop

* fix code style

* fix COVERAGE_FILE path

test=develop

* remove coverage package

test=develop

* test = develop, run coverage as module

27fb9cad

supports collective communicated training (#18175) · b7128bac

由 HaoRen 提交于 6月 27, 2019

* fix prepare context redundant code problem, optimize executor by caching create_varaiables
test=develop

* supports collective training in executor

* make fetch_list runable with variables, add more unittest for use_program_cache
test=develop

* fix comment
test=develop

* use unique name for nccl_id

* supports output to stream in program_to_code

* insert sync_comm_stream before regularization; add skip_op_callstack capability in program_to_code

* set op role in collective training

* add collective op role

* remove orig file

* add build optimizer by strategy

* add collective strategy

* refine collective strategy

* add multi-process role maker

* refine strategy building factory so that we can easily plugin more strategy

* scale loss grad in collective sgd transpiler

* add support for distributed fc

* code format

* revert some features for dist fc

* add support for distributed fc training

* fix prepare context redundant code problem, optimize executor by caching create_varaiables
test=develop

* supports collective training in executor

* make fetch_list runable with variables, add more unittest for use_program_cache
test=develop

* use unique name for nccl_id

* supports output to stream in program_to_code

* insert sync_comm_stream before regularization; add skip_op_callstack capability in program_to_code

* set op role in collective training

* add collective op role

* fix comment
test=develop

* remove orig file

* add build optimizer by strategy

* add collective strategy

* refine collective strategy

* add multi-process role maker

* refine strategy building factory so that we can easily plugin more strategy

* scale loss grad in collective sgd transpiler

* add support for distributed fc

* code format

* revert some features for dist fc

* add support for distributed fc training

* test=develop
add collective op unittest standard

* test=develop
remove the test_collective directory

* test=develop
remove the test_collective directory

* remove slicegather test

* code format for reducescatter

* update attr of shard_index_op

* Modify macro nccl_helper

* remove test without distribute

* macro collective_helper

* marcro update

* test=develop
update support python3.5

* test=develop change gpu memory use to 0.1 when test

* test=develop
update ut equal func

* test=develop
set flags to 1.5

* test=develop fix pickle dumple  py35

* test=develop
fix divide in slice and add sync_comm_stream
update atol and rtol to 1e-05
rm shard_index op and test
modify read input from file to read from memory
remove origin_program in framework and add i/o in c_sync_calc_stream

* test=develop update unittest sync operator I/O

b7128bac

26 6月, 2019 1 次提交
- H
  
  add ut for pipeline training (#18289) · e42057cd
  由 hutuxian 提交于 6月 26, 2019
  
  e42057cd

Crayon鑫 / Paddle 与 Fork 源项目一致

Crayon鑫 / Paddle
与 Fork 源项目一致