提交 · 08fa98f7cc75187d7a2030e07ae1e349979f27cc · BaiXuePrincess / Paddle

01 8月, 2019 3 次提交

Add the op of unique_with_counts, expand count function of the op unique (#18720) · 3ab1866c

由 wawltor 提交于 8月 01, 2019

* test=develop
Add the op of unique_with_counts, the op is calc the unqiue input of data, and output the corresponding indices and count of data.

* test=develop
Check the input and dtype in the op of unique_with_counts

* test=develop
test=document_preview
update the API.spec for `unique_with_counts`, at the same time, optimize the python api in the op of `unique_with_count`

* test=develop
test=document_preview
Fix some python api problem in the op of `unique_with_counts`, and change the error messsage in this op.

* Fix some API problem in the op of `unique_with_counts`
test=develop
test=document_preview

* test=develop
test=document_preview
Fix the api sample of op `unique_with_counts`, and update api.spec

3ab1866c

L
Fix depthwise conv gpu kernel bug (#18582) · 22fa4c2d
由 LielinJiang 提交于 8月 01, 2019
```
* fix depthwise conv gpu kernel bug, test=develop
* add more depthwise conv test, test=develop
```
22fa4c2d
W
Fix unitest of light nas. (#18931) · c92b78b0
由 whs 提交于 8月 01, 2019
```
test=develop
```
c92b78b0

31 7月, 2019 5 次提交

set fleet_send_batch_num a default value according to trainer num · 233746d8

由 jiaqi 提交于 7月 31, 2019

(1) set fleet_send_batch_num a default value according to trainer num， the previous 80000 is fixed，if trainer num is much less or larger than 100，global shuffle may have timeout error.

(2) fix load one table bug, add barrier

233746d8

C
[DyGraph] Make multi-card program faster (#18892) · 20859c08
由 chengduo 提交于 7月 31, 2019
```
* update parallel.py
test=develop
```
20859c08

Add center Loss Op Support (#18681) · 24f85431

由 HaoRen 提交于 7月 31, 2019

* support center loss
* change tensor copy  api to high level api tensorcopy

* test=develop rewrite the center_loss cuda_kernel to make it faster
and add document of the center loss api,also update test function

* test=document_preview test=develop
update document of center loss

* test=document_preview test=develop
modify API.spec modify test code remove nouse const_cast

24f85431

L
replace paper link (#18861) · d21c3914
由 lvmengsi 提交于 7月 31, 2019
```
Update conv2d transpose link
```
d21c3914
D
make dist unit test exclusive run (#18865) · 2bb296df
由 Dong Daxiang 提交于 7月 31, 2019
```
make dist unit test exclusive run
```
2bb296df

30 7月, 2019 3 次提交
- W
  Make lod_append support variable lod. (#18908) · 6cccab92
  由 whs 提交于 7月 30, 2019
```
test=develop
```
  6cccab92
- D
  
  Add elementwise_pow_op backward implementation and the unit test codes of it. (#18848) · e0a2d4df
  由 danleifeng 提交于 7月 30, 2019
  
  e0a2d4df
- C
  add CPUInplaceTestWithFuseOptimizationOps (#18867) · ecd2bdad
  由 chengduo 提交于 7月 30, 2019
```
test=develop
```
  ecd2bdad
29 7月, 2019 2 次提交

Remove legacy C++ memory optimization codes (#18834) · 8008ab4e

由 Zeng Jinle 提交于 7月 29, 2019

* remove legacy memory optimization codes, test=develop

* follow huihuang's comments,test=develop

* follow luotao's comments, test=develop

8008ab4e

add clear_model interface in fleetwrapper (#18815) · 52c1431e

由 Thunderbrook 提交于 7月 29, 2019

* dump slot

* test

* proto

* dump slot

* test

* proto

* code style

* code style

* code style

* style

* add delete after unseen days

* add unseen days

* code style

* conflict solve
test=develop

* add clear model

* code style
test=develop

* code style
test=develop

52c1431e

28 7月, 2019 2 次提交
- Z
  
  fix affine_channel no_need buffer bug, test=develop (#18844) · 9a8a7a1d
  由 Zeng Jinle 提交于 7月 28, 2019
  
  9a8a7a1d
- L
  Fix drop deconv (#18813) · 829ef262
  由 lvmengsi 提交于 7月 28, 2019
```
* replace link

* update api.spec

* fix mistake
```
  829ef262
27 7月, 2019 2 次提交
- C
  Open fuse optimization ops (#18741) · 4140fe11
  由 chengduo 提交于 7月 27, 2019
```
* open fuse optimization ops
test=develop
```
  4140fe11
- C
  add warning info for CPU_NUM (#18840) · 582cc297
  由 chengduo 提交于 7月 26, 2019
```
test=develop
```
  582cc297
26 7月, 2019 2 次提交

A

Add LeakyReLU MKLDNN support (#18762) · ee022279
由 Adam 提交于 7月 26, 2019

ee022279

Feature/mem opt pass refactor (#18735) · a802da65

由 Zeng Jinle 提交于 7月 26, 2019

* first version memory optimize pass, test=develop

* remove move_tensor_sharing_pass, test=develop

* refine code comments, add unittests, test=develop

* turn off memory_optimize by default, test=develop

* follow huihuang's comments, test=develop

* follow chengduoZH's comments, test=develop

* fix grammar error, add const qualifier, fix pass_test exception message, test=develop

* follow chengduoZH's comments 2nd, test=develop

a802da65

25 7月, 2019 4 次提交

石

Fix examples of API (#18092) · 9dbb62ee

由石晓伟提交于 7月 25, 2019

* fix logical APIs

test=develop

test=document_preview

* fix isfinite

* update matmul comments

* update API.spec

test=document_preview

test=develop

* update API.spec

test=document_preview

test=develop

* update API.spec

test=document_preview

test=develop

9dbb62ee

G
refine launch_ps and role_maker (#18795) · 30562e37
由 guru4elephant 提交于 7月 25, 2019
```
refine launch_ps and role_maker
```
30562e37

Fix shrink-dense and add scale-datanorm (#18746) · c167a4b4

由 fuyinno4 提交于 7月 25, 2019

Fix FleetWrapper:
1. fix shrink dense: just scale show
2. add datanorm scale: divide datanorm's gradient by batch_size

c167a4b4

G
split test_dist_se_resnext.py into 4 testcases (#18743) · 2efb282c
由 guru4elephant 提交于 7月 25, 2019
```
* split test_dist_se_resnext.py into 4 testcases
```
2efb282c

24 7月, 2019 5 次提交

Extend Matmul to support matrix multiplication with multiple heads (#18570) · 220eef60

由 Bob Zhu 提交于 7月 24, 2019

* extend matmul op to support multiple head multiplication

With the support of multiple head, the multiplication of two big matrixes is
split into multiplication of several (head_number) small matrixes. e.g. if
Mat A is [3, 24] and Mat B is [24, 4], when multiple A and B with head_number
as 4, Mat A will be split as 4 matrix of [3, 6] and Mat B will be 4 matrix of
[6, 4]. The result of final matrix will be 4 matrix of [3, 4], i.e. [3, 16].

220eef60

Add python API for appending LoD level (#18702) · 075e1cf7

由 whs 提交于 7月 24, 2019

* Make lod reset op support for append lod level.

* Fix API.spec
test=develop

* Fix unitest.
test=develop

* Add python api for lod append.
test=develop

* Fix API.spec
test=develop

* Fix format of doc.
test=develop

* Fix unitest.
test=develop

* Fix doc.
test=develop

075e1cf7

C
Enhance backward process (#18700) · 8259f141
由 chengduo 提交于 7月 24, 2019
```
* prun backward ops
test=develop
```
8259f141

Modify auc doc. Add output variable description, previously was the scalar... · 25c9b57b

由 JesseyXujin 提交于 7月 24, 2019

Modify auc doc. Add output variable description, previously was the scalar type, now changed to the tuple type.test=develop (#18771)

25c9b57b

add slot to sparse table (#18686) · d8396281

由 Thunderbrook 提交于 7月 24, 2019

The change includes 2 things:

1. save delta model and shrink table are control by the same parameter before, now add delete_after_unseen_days to control shrink table.
2. value in sparse table has no slot before, now add slot in sparse table, and add DownpureCtrAccessor to support the new meta.
test=develop

d8396281

23 7月, 2019 3 次提交

support patch data, add load_one_table, fix bug (#18509) · d18aabb4

由 jiaqi 提交于 7月 23, 2019

（1）support patch data （merge slots of instances of same line id, modify dense layer which
changes its size）
（2）add fleet load_one_table interface, support load from paddle model and load from pslib model
（3）fix push sparse bug which cause push sparse cost more time（about 10% in my testcase）
（4）when some slots are not in one of your network (join/update, etc.)，data feed、collect label info、push/pull sparse will skip these slots， instead of throw error.
（5）add more debug info in TrainFilesWithProfiler

d18aabb4

C
Make fuse_optimizer_op_pass also work when the model contains sparse gradients. (#18664) · fd3aad6c
由 chengduo 提交于 7月 23, 2019
```
* support sparse gradients
test=develop
```
fd3aad6c

supports distributed classification (#18690) · 157211c4

由 Yi Liu 提交于 7月 23, 2019

* supports distributed classification training
* update API.spec
* fix evenly division in python3
* change "index_range" to "index_num" in shard_index operator
test=document_preview
test=develop

157211c4

22 7月, 2019 5 次提交

Z

add more traceback to py_reader error msg, test=develop (#18722) · d07ad4c6
由 Zeng Jinle 提交于 7月 22, 2019

d07ad4c6

Fix random test_recurrent_op failure (#18718) · a3028bb7

由 Huihuang Zheng 提交于 7月 22, 2019

The change includes 3 things:

1. Set CPU_NUM to 1 in the tests because the ParallelExecutor will print warning that CPU_NUM is not set and use default 1.

2. Old tests compare two RNNs, hand written simple RNN and same RNN built by Paddle, but initialized RNN weights in numpy random and Paddle random separately. Fixed it by setting weights and bias values.

3. Also set numpy random seed in the tests. Now the two RNNs diff can be smaller (rtol from 0.1, 0.2 to. 0.01) in the tests.

test=develop

a3028bb7

T
Revert "Add LeakyRelu MKLDNN support (#18656)" (#18723) · bd22453f
由 Tao Luo 提交于 7月 22, 2019
```
test=develop
```
bd22453f
T
do some odd jobs (#18641) · d8458483
由 tangwei12 提交于 7月 22, 2019
```
do some odd jobs, test=develop
```
d8458483
G
split different comm method for mnist distributed training (#18715) · ebf9797e
由 guru4elephant 提交于 7月 22, 2019
```
* split different comm method for mnist distributed training
```
ebf9797e

19 7月, 2019 3 次提交

Support memory eager deletion on recurrent OP (#17710) · 89bc3fd8

由 Huihuang Zheng 提交于 7月 19, 2019

Test PaddingRNN on V100 GPU device.

Test configuration: large model, padding mode (which is the mode using recurrentOp), one GPU.

GPU memory (MiB): 6414 (this PR) vs 6837 (without this PR)
Speed (steps/s): 10.28 (this PR) vs 9.89 (without this PR)

89bc3fd8

A
Add LeakyRelu MKLDNN support (#18656) · d6b6a337
由 Adam 提交于 7月 19, 2019
```
test=develop
```
d6b6a337
T
add check of executor (#17986) · 0b9acb49
由 tangwei12 提交于 7月 19, 2019
```
* add check of executor, test=develop
```
0b9acb49

18 7月, 2019 1 次提交

Feature/auto_growth_allocator (#18561) · ae58afc5

由 Zeng Jinle 提交于 7月 18, 2019

* feature/auto_growth_allocator, test=develop

* add unittest of AlignedAllocator, test=develop

* try to turn on auto_growth to test on CI, test=develop

* fix segmentation fault in mixed_vector.h, test=develop

* add unittests, test=develop

ae58afc5

BaiXuePrincess / Paddle 与 Fork 源项目一致

BaiXuePrincess / Paddle
与 Fork 源项目一致