提交 · 51b081230aecbec3c6614713a82ba6fdf74f6c35 · BaiXuePrincess / Paddle

22 11月, 2022 1 次提交

[remove fluid] under fleet meta_optimizers_wz (#47888) · 51b08123

由 wangzhen38 提交于 11月 22, 2022

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* update

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

* [remove fluid] under fleet meta_optimizers_wz

51b08123

03 11月, 2022 1 次提交

[CodeStyle][py2][U008] remove unnecessary args in `super()` (#47549) · 3de3e45e

由 Nyakku Shigure 提交于 11月 03, 2022

* [CodeStyle][py2][U008] remove unnecessary args in `super()`

* remove remained args

* revert changes in test_pylayer_op

* Revert "revert changes in test_pylayer_op"

This reverts commit ff185a9ae738afac3b0264f61bde6c6b7f72e7c4.

* revert some changes in example code

3de3e45e

01 11月, 2022 2 次提交

[CodeStyle][E711] use `is`/`is not` for comparison with `None` (#47452) · a35a4a53

由 Nyakku Shigure 提交于 11月 01, 2022

* [CodeStyle][E711] use `is`/`is not` for comparison with `None`

* `self.assertTrue($A is None)` -> `self.assertIsNone($A)`

* `self.assertTrue($A is not None)` -> `self.assertIsNotNone($A)`

* `self.assertFalse($A is None)` -> `self.assertIsNotNone($A)`

* `self.assertEqual($A, None)` -> `self.assertIsNone($A)`

* `self.assertNotEqual($A, None)` -> `self.assertIsNotNone($A)`

a35a4a53

[CodeStyle][E712] use `if cond`/`if cond is True` for comparison with `True` (#47464) · 5a2ab683

由 Nyakku Shigure 提交于 11月 01, 2022

* [CodeStyle][E712] use `if cond`/`if cond is True` for comparison with `True`

* revert changes in fluid

* revert unrelated file

* revert changes in norm

* revert changes in auto_parallel_amp

* fix norm and auto_parallel_amp

* revert a typo fix due to fixed at #47477

5a2ab683

23 10月, 2022 1 次提交
- N
  [CodeStyle][black] use black instead of yapf (#46014) · 7097630f
  由 Nyakku Shigure 提交于 10月 23, 2022
```
* update config

* re-blacken python code

* temporarily disable date and diff_py_file

* skip a format
```
  7097630f
09 6月, 2022 1 次提交

Add nproc_per_node for DistributedFusedLamb (#43295) · 6678def9

由 sneaxiy 提交于 6月 09, 2022

* add nproc_per_node for DistributedFusedLamb

* fix nproc_per_node communicator bug

* fix ring_id = 1 init bug

* fix ci

* fix test_parallel_executor_mnist.py

6678def9

05 6月, 2022 1 次提交

【code format check upgrade】 step2：yapf (#42944) · a072fca8

由 Sing_chan 提交于 6月 05, 2022

* use yapf to format all python file

* yapf exclude two unittests file for they rely on writing and reading file, and format will break them

* disable diff_py_file because too many diff files cause command following failed

a072fca8

08 9月, 2021 1 次提交

Enable program passes on Fleet APIs (#34955) · 5f369881

由 Zeng Jinle 提交于 9月 08, 2021

* add fleet api for program pass

* turn on apply pass for CI test

* fix disable fuse_all_optimizer bug

* try to test ci

* fix CI

* fill unspecified op role

* fix fuse_allreduce

* add ut to improve coverage

* remove useless change

* improve c++ coverage

* follow some comments

* test ir pass pipeline

* update doc

* reduce ut time again

5f369881

01 7月, 2021 1 次提交
- Y
  
  gradient scale (#33862) · 57aabbab
  由 Yuang Liu 提交于 7月 01, 2021
  
  57aabbab
13 5月, 2021 1 次提交
- W
  
  fix wait server ready (#32889) · a8625aaf
  由 WangXi 提交于 5月 13, 2021
  
  a8625aaf
06 5月, 2021 1 次提交
- Z
  
  update 2.0 public api in distributed (#32695) · 70eb435c
  由 zhiboniu 提交于 5月 06, 2021
  
  70eb435c
07 4月, 2021 1 次提交

【NPU】Merge ascend GE&distributed code by 0208 from ascendrc (#31957) · 8c7c53b3

由 zhang wenhui 提交于 4月 07, 2021

* Ascend rc (#30483)

* Fix compilcation on CANN20.1 and older (#30494)

Fix compilcation on CANN20.1 and older

* Add distribution supported (#30578)

Add distribution supported

* Build praser for Hcom* operators (#30627)

Build praser for Hcom* operators

* Pass device_ids info from launch to trainer. (#30632)

Pass device_ids info from launch to trainer

* Add Hccl program group (#30642)

Add Hccl program group

* Add startup bash files of test_ascend_group. (#30645)

Add startup bash files of test_ascend_group

* cleanup (#30646)

cleanup test_ascend_group.py

* [Feature] Build parser to support distributed training (#30658)

[Feature] Build parser to support distributed training

* fix compilation on ascend-20.1 (#30722)

fix compilation on ascend-20.1

* Dev/fix ascend string (#30749)

Dev/fix ascend string

* code style (#30781)

code style

* Merge ascend_optimizer and ascend_parser. (#30776)

Merge ascend_optimizer and ascend_parser.

* Ascendrc add converted op : [range/equal/range/uniform_random/expand/squeeze], fix cast op bug  (#30797)

Ascendrc add converted op : [range/equal/range/uniform_random/expand/squeeze], fix cast op bug

* Add paddle ascend distribution training supported (#30796)

Add paddle ascend distribution training supported

* pass cxx_flags to gloo cmake (#30857)

* Destroy session first. (#30954)

Destroy session first.

* merge

* fix, test=develop

* fix, test=develop

* fix style, test=develop

* fix, test=develop

* fix

* fix log fatal, test=develop

* fix enforce style, test=develop

* fix, test=develop

* fix, test=develop

* fix rccl, test=develop

* fix test, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix node_num, test=develop

* fix ids str, test=develop

* fix ids str, test=develop

* fix ids str, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix style code, test=develop

* fix style code, test=develop

* fix style code, test=develop

* fix style code, test=develop
Co-authored-by: Nhutuxian <hutuxian2011@sina.cn>
Co-authored-by: Ngongweibao <weibao.gong@gmail.com>
Co-authored-by: NVoid Main <voidmain1313113@gmail.com>
Co-authored-by: NLeo Chen <chenqiuliang@baidu.com>
Co-authored-by: Ndingsiyu <18369187719@163.com>
Co-authored-by: NOleNet <olenet@126.com>

8c7c53b3

05 2月, 2021 1 次提交
- L
  
  [Kunlun] add gen_bkcl_id_op, support multi XPU cards training using multiprocess (#30858) · 4a8b8b45
  由 liuyuhui 提交于 2月 05, 2021
  
  4a8b8b45
01 2月, 2021 1 次提交
- W
  
  Fleet distributed strategy support pure fp16 (#30754) · 31ed9c9e
  由 WangXi 提交于 2月 01, 2021
  
  31ed9c9e
17 12月, 2020 1 次提交
- W
  
  fleet sync build strategy, test=develop (#29732) · 9cbcc6ca
  由 WangXi 提交于 12月 17, 2020
  
  9cbcc6ca
26 11月, 2020 1 次提交
- W
  
  Fix multi nccl comm & wait server ready (#28663) · e931c7ba
  由 WangXi 提交于 11月 26, 2020
  
  e931c7ba
20 9月, 2020 1 次提交

【paddle.fleet】Fix/role maker api fix (#27326) · d6b54de4

由 tangwei12 提交于 9月 20, 2020

* fix fleet util and gloo

* fix worker endpoints

* fix

* fix UT

* fix gloo

* fix gloo

* update gloo

* update gloo

* update gloo

* update gloo

* update gloo

* fix gloo wrapper for hdfs

* add file gloo and UT

* fix UT

* fix UT

* fix UT

* hide public method of RoleMaker

* fix UT

* GPU fleetrun support gloo

* parameterserver fleetrun support gloo

* add UT

* add UT

* fix UT

* fix get server endpoint

* fix get server endpoint

* fix UT

* hide public method of rolemaker

* hide public method of rolemaker

* hide public method of rolemaker

* Update test_fleet_rolemaker_new.py

* hide public method of rolemaker

* hide public method of rolemaker

d6b54de4

10 9月, 2020 1 次提交
- 1
  【paddle.fleet】parameter_server_optimizer support auto_strategy (#27181) · 60c3ef3a
  由 123malin 提交于 9月 10, 2020
```
* parameter_server_optimizer support auto_strategy
```
  60c3ef3a
07 9月, 2020 1 次提交
- D
  【paddle.fleet】add auto parallel L1 implementations (#27090) · 0443b480
  由 Dong Daxiang 提交于 9月 07, 2020
```
* add auto parallel L1 implementation
test=develop
```
  0443b480
21 8月, 2020 1 次提交
- D
  【paddle.fleet】Meta from optimizer (#26392) · 83cd1859
  由 Dong Daxiang 提交于 8月 21, 2020
```
* consider the combination of different strategies to work together
```
  83cd1859
18 8月, 2020 1 次提交
- M
  add feature to fleet2.0 role_maker, distribute_strategy, test=develop (#26267) · cd48bdad
  由 mapingshuo 提交于 8月 18, 2020
```
* add feature to fleet2.0 role_maker, distribute_strategy, test=develop
```
  cd48bdad
13 8月, 2020 1 次提交
- D
  【paddle.fleet】paddle.fleet -> paddle.distributed.fleet. (#26186) · 50a5bcfc
  由 Dong Daxiang 提交于 8月 13, 2020
```
* move paddle.fleet to paddle.distributed.fleet
```
  50a5bcfc
11 8月, 2020 1 次提交

Paddle-2.0 API directory migration (#25898) · 2efcb481

由 pangyoki 提交于 8月 10, 2020

* Directory migration, test=develop

* Change imperative from paddle init to paddle framework, test=develop

* Fixed jit bug, test=develop

* default static mode, test=develop

* fixed format and create parameter belongs to framework, test=develop

* Fixed import package, test=develop

* fix __init__ format, test=develop

* fixed alias problem

* fixed paddle.enable_imperative problems, test=develop

* Add unittest

* delete install_check comment

* Fixed unittest timeout

* fixed unittest error

* move Program default_xx_program to static package

* optimize unittest method

* fixed framework __init__ format

* fixed jit path

* delete alias

* move jit to paddle

* Fixed unittest format

* fixed paddle.default_main_program

* Fixed save load API in paddle __init__.py

* fixed ci paddle.imperative.to_variable

2efcb481

03 8月, 2020 1 次提交

【paddle.fleet】Fleet run graph in Executor and add two more strategies (#25844) · 8d2896f1

由 Dong Daxiang 提交于 8月 03, 2020

* split meta optimizer files
* add graph execution in execution, update two properties in DistributedStrategy, unit tests for these features

8d2896f1

30 7月, 2020 1 次提交

Integrated Trainer of Parameter Server (API add... · caa90a65

由 tangwei12 提交于 7月 30, 2020

Integrated Trainer of Parameter Server (API add `fluid.contrib.layers.sparse_embedding` only) (#22957)

* Integrated Trainer of Parameter Server

caa90a65

29 7月, 2020 1 次提交
- D
  Generate final strategy (#25782) · a96d54ac
  由 Dong Daxiang 提交于 7月 29, 2020
```
* refine strategy compiler and meta optimizers
make async as a_sync
```
  a96d54ac
28 7月, 2020 1 次提交

add more settings for distributed strategy (#25685) · 920d998f

由 Dong Daxiang 提交于 7月 28, 2020

* add more settings for distributed strategy
Basically, DistributedStrategy has several parts of configurations:
- BuildStrategy: the same as paddle.fluid.BuildStrategy, but the distributed arguments are moved out of BuildStrategy
- ExecutionStrategy: the same as paddle.fluid.ExecutionStrategy
- collective communication configs: nccl_comm_num, hierarchical allreduce and so on
- distributed algorithms: async_update(mainly used in PS), lars, lamb and so on

920d998f

23 7月, 2020 1 次提交
- D
  fix gen nccl id bug (#25669) · 28064c2d
  由 Dong Daxiang 提交于 7月 23, 2020
```
* fix gen nccl id bug
```
  28064c2d
20 7月, 2020 1 次提交
- D
  fleet base initial implementation and the API (#25442) · e657d706
  由 Dong Daxiang 提交于 7月 20, 2020
```
refactor fleet api under paddle.fleet
update DistributedStrategy
```
  e657d706

BaiXuePrincess / Paddle 与 Fork 源项目一致

BaiXuePrincess / Paddle
与 Fork 源项目一致