提交 · 6e7bfe30a6e9584a8ce21a6f6ef66a2bfb3b1097 · BaiXuePrincess / Paddle

22 2月, 2020 1 次提交
- T
  SYNC with communicaotor (#22344) · 66a31501
  由 tangwei12 提交于 2月 22, 2020
```
* add sync communicator and implement
```
  66a31501
17 2月, 2020 1 次提交
- 1
  
  support dumping params/grads in transpiler mode (#22490) · 00594c1c
  由 123malin 提交于 2月 17, 2020
  
  00594c1c
16 2月, 2020 2 次提交
- 1
  
  test=develop, add distributed tools (#22623) · e59463ef
  由 123malin 提交于 2月 16, 2020
  
  e59463ef
- T
  add texttable for pretty flag output (#22584) · 1aab3e61
  由 tangwei12 提交于 2月 16, 2020
```
pretty print for communicator flag
```
  1aab3e61
12 2月, 2020 1 次提交
- T
  fix bug with compiledProgram (#22495) · b0675c81
  由 tangwei12 提交于 2月 12, 2020
```
* add thread barrier for the compiled program
```
  b0675c81
11 2月, 2020 1 次提交

multi-loss optimization by adding a DownpourOpt worker (#22025) · 2235ee1a

由 yaoxuefeng 提交于 2月 11, 2020

* update

* update test=develop

* update compile set test=develop

* update compile set test=develop

* update test=develop

* update test=develop

* update test=develop

* update compile setting test=develop

* update compile setting test=develop

* update run demo test=develop

* update test=develop

* update test=develop

* fix test=develop

* update test=develop

* update test=develop

* update test=develop

* update test=develop

* update test=develop

* update test=develop

* update test=develop

* update test=develop

* update test=develop

* update format test=develop

* update format test=develop

* update style test=develop

* update style test=develop

* change style test=develop

* change style test=develop

* change style test=develop

* add dataset unittest test=develop

* update test=develop

* update for record test=develop

* udpate style for record test=develop

* update for record test=develop

* update for record test=develop

* update for record test=develop

* fix format test=develop

* update test=develop

* update test=develop

* update test=develop

* update test=develop

* update test=develop

2235ee1a

05 2月, 2020 1 次提交
- X
  add hdfs ls retry time and sleep time, fix save inference (#22433) · 6e4f39a0
  由 xujiaqi01 提交于 2月 05, 2020
```
* add hdfs ls retry time and sleep time, fix save inference
* test=develop
```
  6e4f39a0
03 2月, 2020 1 次提交
- T
  fix bug with half (#22378) · 7e2665c5
  由 tangwei12 提交于 2月 03, 2020
```
* fix bug with half communicator
```
  7e2665c5
02 2月, 2020 1 次提交
- X
  add GeneralRoleMaker (#22295) · 371f377b
  由 xujiaqi01 提交于 2月 02, 2020
```
* add GeneralRoleMaker which is for general usage
* test=develop
```
  371f377b
17 1月, 2020 1 次提交
- T
  integrated HALF_ASYNC to communicator (#21869) · 82bc814a
  由 tangwei12 提交于 1月 17, 2020
```
* add half_async in the communicator
* fix DistributedStrategy
```
  82bc814a
06 1月, 2020 1 次提交
- 1
  add distributed_strategy (#21710) · 7fb817d4
  由 123malin 提交于 1月 06, 2020
```
* add distributed_strategy
```
  7fb817d4
31 12月, 2019 1 次提交
- W
  
  fix sync_batch_norm hang in fleet (#21838) · 3ec289a6
  由 WangXi 提交于 12月 31, 2019
  
  3ec289a6
05 12月, 2019 1 次提交
- L
  
  bugfix: construct a DistributedStrategy instance if the passed one is None (#21545) · da75ac8b
  由 lilong12 提交于 12月 05, 2019
  
  da75ac8b
28 11月, 2019 1 次提交
- X
  fix fleet save bug (#21362) · f1178e9d
  由 xujiaqi01 提交于 11月 28, 2019
```
* fix fleet save bug of save_infernece_model
* test=develop
```
  f1178e9d
26 11月, 2019 1 次提交
- Z
  Fix some typos in AMP. (#21354) · be2e3e67
  由 Zhen Wang 提交于 11月 26, 2019
```
* fix some typos in AMP. test=develop

* delete useless codes. test=develop
```
  be2e3e67
25 11月, 2019 1 次提交
- T
  print table stat info for pslib (#21296) · 9a7832f8
  由 Thunderbrook 提交于 11月 25, 2019
```
* print table stat
test=develop

* notes
test=develop

* notes
test=develop
```
  9a7832f8
21 11月, 2019 3 次提交

fix fs_client_param bug (#21212) · 319d2ba9

由 xujiaqi01 提交于 11月 21, 2019

* fix fs_client_param bug， user can set this config through fleet_desc_file or fleet config
* test=develop

319d2ba9

solve pslib core in stop worker (#21263) · 0d17c1b8

由 Thunderbrook 提交于 11月 21, 2019

* general table

* add sparse table
test=develop

* no cvm
test=develop

* add no_cvm
test=develop

* add note
test=develop

* code style
test=develop

* code style
test=develop

* code style
test=develop

* code style
test=develop

* code style
test=develop

* add key of optimizer
test=develop

* solve pslib stop core
test=develop

* barrier
test=develop

* add notes
test=develop

0d17c1b8

X
fix fleet util bug (#21254) · eca66f31
由 xujiaqi01 提交于 11月 21, 2019
```
* fix fleet util bug in save paddle inference model
* test=develop
```
eca66f31

20 11月, 2019 2 次提交

support general embedding params (#21217) · 349e82d6

由 Thunderbrook 提交于 11月 20, 2019

* general table

* add sparse table
test=develop

* no cvm
test=develop

* add no_cvm
test=develop

* add note
test=develop

* code style
test=develop

* code style
test=develop

* code style
test=develop

* code style
test=develop

* code style
test=develop

* add key of optimizer
test=develop

349e82d6

D
update worker_num for MPISymetricRoleMaker (#20798) · ccbdd7aa
由 Dong Daxiang 提交于 11月 20, 2019
```
test=develop
```
ccbdd7aa

15 11月, 2019 2 次提交

X
fix cache table bug, add save_paddle_inference_model, fix hdfs util bug (#21052) · 23876de5
由 xujiaqi01 提交于 11月 15, 2019
```
* fix cache table bug
* add save_paddle_inference_model
* fix hdfs util bug
* test=develop
```
23876de5

add copy table (#21086) · 9e045170

由 xujiaqi01 提交于 11月 15, 2019

* copy some feasigns and corresponding embeddings from one sparse table to another
* copy all feasigns and corresponding embeddings from one sparse table to another
* copy all dense params from one table to another
* copy some local vars to other local vars

9e045170

12 11月, 2019 1 次提交

modify the implementation of save_persistables and save_inference_model for... · 53148e06

由 lilong12 提交于 11月 12, 2019

modify the implementation of save_persistables and save_inference_model for fleet collective mode (#20802)

* modify the implementation of  save_persistables and save_inference_model functions for fleet collective, test=develop

* add ut, test=develop

53148e06

04 11月, 2019 1 次提交
- T
  find lookup table in order (#20932) · 5970e8ac
  由 Thunderbrook 提交于 11月 04, 2019
```
test=develop
```
  5970e8ac
31 10月, 2019 3 次提交

Fix Paddle Cloud role maker (#20860) · 16596f64

由 Chengmo 提交于 10月 31, 2019

* fix PaddleCloud Role maker & add warning in distribute transpiler  & change rpc_retry_times

16596f64

B

fix hdfs.download, test=develop (#20907) · ac87d4e6
由 Bai Yifan 提交于 10月 31, 2019

ac87d4e6

support dump param of model into afs (#20302) · 59bcdc8a

由 Thunderbrook 提交于 10月 31, 2019

* support dump param to afs
test=develop

* code style
test=develop

* code style
test=develop

* dump param
test=develop

* dump param
test=develop

* dump param
test=develop

* dump param
test=develop

59bcdc8a

25 10月, 2019 1 次提交

fix several sparse table issuses (#20686) · 48669aa8

由 xujiaqi01 提交于 10月 25, 2019

* no longer need to define all embedding layers (no one less) of all slots in each program. make trainer_param repeated in ps.proto.
* add find_distributed_lookup_table_grads instead of hard code GRAD
* support embedding stop gradient. push sparse has error before fix this.* 
* fix fill sparse, skip slots which do not have embedding. each slot's embedding in a sparse table should be used in all training programs before fix this.
* fix pull sparse, skip slots which do not have embedding.
* fix collect feasign label info, skip slots which do not have embedding.
* support when there are multi sparse tables in one or multi training programs, each program can pull/push its own related sparse tables instead of all sparse tables.
* test=develop

48669aa8

18 10月, 2019 1 次提交
- X
  add check nan / inf in downpour worker (#20694) · 5223b0dd
  由 xujiaqi01 提交于 10月 18, 2019
```
* add check nan / inf in downpour worker during training
* test=develop
```
  5223b0dd
15 10月, 2019 3 次提交

Fix communicator slow bug & fix communicator stop bug (#20366) · 940c6ff1

由 Chengmo 提交于 10月 15, 2019

* test=develop,Fix communicator slow bug

* test=develop, delete if() in stop_worker()

* test=develop

* fix UT, test=develop

* fix bug in fetch handler, test=develop

* fix bug in fetch handler, test=develop

* test=develop, fix fetch barrier bug

* test=develop, bug fix

* test=develop, bug fix

* test=develop, fix bug

940c6ff1

W

fix dgc test and bug when not set trainers_endpoints_, test=develop (#20617) · cadc6a97
由 WangXi 提交于 10月 14, 2019

cadc6a97
M
Fleet: deal with special case: strategy is None (#20359) · f55d1c68
由 mapingshuo 提交于 10月 15, 2019
```
* special case: strategy is None
```
f55d1c68

14 10月, 2019 1 次提交
- T
  dump fix dov vec file num (#20539) · f76a32df
  由 Thunderbrook 提交于 10月 14, 2019
```
* support dump multi file
test=develop

* dump fix num file
test=develop
```
  f76a32df
12 10月, 2019 1 次提交
- Z
  
  fix converter , test=develop (#20522) · b5219920
  由 zhang wenhui 提交于 10月 12, 2019
  
  b5219920
11 10月, 2019 1 次提交
- Z
  fix pslib datanorm double bug (#20297) · b82e6520
  由 zhang wenhui 提交于 10月 11, 2019
```
* fix fc sort . test=develop
```
  b82e6520
07 10月, 2019 1 次提交
- Z
  
  fix fleet_desc delete_after_unseen_day bug in node.py (#20091) · b28d4a82
  由 zhang wenhui 提交于 10月 07, 2019
  
  b28d4a82
30 9月, 2019 1 次提交
- C
  Add GEO-SGD distribute training algorithm (#20018) · 728ec1b4
  由 Chengmo 提交于 9月 30, 2019
```
* refector geo sgd & communicator
```
  728ec1b4
24 9月, 2019 1 次提交

support change shuffle and train thread num (#19841) · cedc0477

由 xujiaqi01 提交于 9月 24, 2019

* support change shuffle thread num
* support change train thread num
* fix receive shuffle data of each channel
* data norm stop gradient
* add check thread_tensor type and root_tensor type when merge metric
* remove sleep in shuffle, add config
* add config of pslib client to client communication
* fix xbox str
* add data norm op testcase
* add flush in trainer finalize

cedc0477

23 9月, 2019 1 次提交

Forward recompute3 (#19913) · 9901f696

由 mapingshuo 提交于 9月 23, 2019

* add recompute based checkpoints methods for large batch training
test=develop

* add append_backward_with_forward_recomputation
test=develop

* refine optimizer
test=develop

* update backward and optimizer
test=develop

* make Variable usable
test=develop

* add recompute code

* refine optimizer
test=develop

* refine addup _append_backward_ops_with_checkpoints_
1) for recompute part, just cache the grad_op_desc without appending to block
2) before appending grad_op_desc to backward part, addup_repetitive_vars, remove unused branch
test=develop

* make method private

* add recompute strategy into DistributedStrategy
test=develop

* checkpoint version3
test=develop

* remove some print information
test=develop

* remove unused sumop
test=develop

* try to fix recompute with graph building modules

* add input names to vars should be held

* add memory debug tool

* backup backward

* Fix bugs

* add backward desc for op not in any segments

* add exception info for sub_block

test=develop

* modify code style

test=develop

* modify code style

test=develop

* remove print functions

test=develop

* add API spec

test=develop
test=document_preview

* make Recompute a child class of Optimizer

test=develop
test=document_preview

* add API spec

test=develop
test=document_preview

* modify API spec

test=develop
test=document_preview

* add document for Recompute

test=develop
test=document_preview

* change API doc of Rcompute

test=develop
test=document_preview

* code cleaning

test=develop
test=document_preview

* modify API spec

* fix bugs when segments hold no element

* add testcase for Recompute Optimizer

test=develop
test=document_preview

* add test for apply_gradient, and code cleaning

test=develop
test=document_preview

* add test case for load function

* enable CI

test=develop
test=document

* add test case

test=develop
test=document_preview

* add sample code for 4 function of recompute optimizer

test=develop
test=document_preview

9901f696

BaiXuePrincess / Paddle 与 Fork 源项目一致

BaiXuePrincess / Paddle
与 Fork 源项目一致