提交 · 2c70b844fc934992a741f7bc76a85bd9cddd126a · 机器未来 / Paddle

16 9月, 2021 2 次提交

L

rename the auto parallel suffix (#35765) · 2c70b844
由 lilong12 提交于 9月 16, 2021

2c70b844

Python support register pass via PassDesc (#35602) · bab39eb2

由 wuhuanzhou 提交于 9月 16, 2021

PR主要功能：针对fusion等子图替换场景，支持Python侧开发并注册Pass。

背景
Pass是指输入一个深度学习计算图Graph，依照一定条件进行修改，输出修改后的Graph的过程；
当前PaddlePadle框架编写Pass代码存在以下问题：
用户需要手写Graph的条件匹配、在Graph上的修改代码；
对Graph操作需要深入底层框架代码，了解Graph的结构，并且知道相关Pass写法；
我们提出了针对fusion等子图替换类Pass的优化方案以支持用户在Python侧开发注册Pass，提升二次开发体验：
用户只需要输入匹配和替换的子图描述，由深度学习框架编写的代码来生成匹配和替换的逻辑，不需要用户对Graph进行匹配和替换操作；
API级别的替换，用户可以通过Paddle的Python API构造子图，从而不需要知道Graph的结构，也能写Paddle的Graph Pass代码

bab39eb2

15 9月, 2021 3 次提交
- W
  add inplace logic into new_executor (#35618) · bd79ae09
  由 wanghuancoder 提交于 9月 15, 2021
```
* add inplace logic into new_executor, test=develop

* check shape and add inplace FLAGS, test=develop

* refine, test=develop

* refine, test=develop
```
  bd79ae09
- 王
  clip op extra information when export model. (#35447) · 4d236354
  由王明冬提交于 9月 15, 2021
```
* clip op extra information when export model,test=ocr

* rename clip_extra parameter to kwargs in save_inference_model, test=ocr
```
  4d236354
- W
  
  [hybrid] out data parallel as optimizer sharding parallel (#35593) · 78465703
  由 WangXi 提交于 9月 15, 2021
  
  78465703
14 9月, 2021 6 次提交
- Y
  
  [HETERPS]fix make error when pscore equals on · 3493c46e
  由 yaoxuefeng 提交于 9月 14, 2021
  
  3493c46e
- J
  Add dropout convert test (#35488) · 6327c33b
  由 JingZhuangzhuang 提交于 9月 14, 2021
```
* add dropout convert test

* modify dropout convert test
Co-authored-by: Nxiaoxiaohehe001 <hiteezsf@163.com>
```
  6327c33b
- W
  
  skip sharelod, test=develop (#35625) · 16e40513
  由 wanghuancoder 提交于 9月 14, 2021
  
  16e40513
- A
  
  Refactor StreamAnalyzer and EventManager from InterpreterCore (#35711) · 8528dd9f
  由 Aurelius84 提交于 9月 14, 2021
  
  8528dd9f
- Y
  
  [hybrid performance] Optimize Pipeline Scheduler (#35680) · 04fdb10a
  由 Yuang Liu 提交于 9月 14, 2021
  
  04fdb10a
- A
  Intergrate StandaloneExecutor in Static.Executor Interface with... · 4bc08530
  由 Aurelius84 提交于 9月 14, 2021
```
Intergrate StandaloneExecutor in Static.Executor Interface with FLAGS_USE_STANDALONE_EXECUTOR (#35628)

* Intergrate StandaloneExecutor in Static.Executor Interface with FLAGS_USE_STANDALONE_EXECUTOR

* Enhance unittest and clean code in StandaloneExecutor

* polish unittest
```
  4bc08530
13 9月, 2021 1 次提交
- L
  
  add lstm qat models scales (#35382) · 1ee237c1
  由 lidanqing 提交于 9月 13, 2021
  
  1ee237c1
12 9月, 2021 1 次提交
- 王
  
  fix some op extra error, test=develop (#35667) · 83424033
  由王明冬提交于 9月 12, 2021
  
  83424033
11 9月, 2021 2 次提交
- W
  refactor gc (#35525) · adaa207b
  由 wanghuancoder 提交于 9月 10, 2021
```
* refactor gc, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop

* gc each tensor, test=develop

* refine, test=develop
```
  adaa207b
- 王
  
  register the with_quant_attr attribute for all operattor. test=develop (#35591) · 8412d6c0
  由王明冬提交于 9月 11, 2021
  
  8412d6c0
10 9月, 2021 2 次提交

add cumprod op (#35185) · 4e509f46

由 hlygit66666 提交于 9月 10, 2021

* add test_cumprod_op

* Revert "add test_cumprod_op"

This reverts commit c96cf6dff5d09ae7d8cc72c1e8ae4369a153aa19.

* recommit

* add error message

* test input(x) initialize

* test use cpu

* update test code

* add test type

* add test case

* solve ci problem

* add complex case test

* add complex case test

* fix review problem

* fix conflict

* fix some docs

* change test case

* change test case

* fix review problems again

* fix docs

* fix inclusivescan bug

4e509f46

import ska flat_hash_map (#34464) · 3d9603dc

由 chentianyu03 提交于 9月 10, 2021

* import ska flat_hash_map

* add define NOMINMAX macro to fix windows build failed bug

* add brackets to std::max in flat_hash_map

* move flat_hash_map directions

* modify namespace to paddle

* modify namespace to paddle

* modify namespace to paddle

* modify namespace to paddle

* rm not used map.h and replace with op_info

3d9603dc

08 9月, 2021 5 次提交

[Auto Parallel] Integrate all modules (#35483) · 12155358

由 Yulong Ao 提交于 9月 08, 2021

* add auto_parallel dir

* mv to paddle.distributed

* add shard_xx api

* add distributed attrs for var

* add ut, test=develop

* add dist

* update

* update

* update

* update

* update

* update, test=develop

* update, test=develop

* update, test=develop

* update, test=develop

* update, test=develop

* update, test=develop

* update, test=develop

* update

* update

* update

* update

* update

* update, test=develop

* update, test=develop

* update

* update

* delete unused proto

* resotre op_desc

* restore type_defs

* update var_desc

* remove dimss_mapping for proto_pybind

* update interface.py

* update framework.py

* update

* update

* add auto_parallel dir

* mv to paddle.distributed

* add shard_xx api

* add distributed attrs for var

* add ut, test=develop

* [WIP] Add the auto completion feature and related codes

* [WIP] Improve the auto completion and related codes

* [WIP] Make the auto completion to support data-parallel

* [WIP] Make the completion support mp and dp+mp

* [WIP] Refactor auto completion unit test for MLP

* [WIP] Refactor the implementation of DistributedOperatorImpl

* [WIP] Improve dims_mapping update rule and fix a bug

* [WIP] Support auto completion for one transformer decoder layer

* [WIP] Add a minor change

* [WIP] Fix a bug within the uint test

* Shard XShape tensor, add embedding completion and refactor code

* Add the distributed_operators dir to setup.py.in

* Improve the completion process and add the unittest for gpt

* fix process_mesh ut

* fix process_mesh ut

* update

* update, test=develop

* Add support for automatically completing distributed attrs of special ops

* update

* update

* update

* fix doc sample codes, test=develop

* improve coverage, test=develop

* add static_mode check, test=develop

* Model the cluster for cost model and physical mapping

* update, test=develop

* add set_placement, test=develop

* Add the check to make sure the candidate tensors' size is great than zero

* update doc, test=develop

* update doc, test=develop

* update doc, test=develop

* update doc, test=develop

* update, test=develop

* Auto mark dist attrs annotated by user

* update ndarray to nested list, test=develop

* update, test=develop

* Add auto-completion module for auto-parallel (based on PR#33804)

* Remove unnecessary files

* Remove unrelated files for the auto completion pr

* Update the unit test to improve the coverage

* Modify codes based on reviews

* Minor changes for CI

* Improve some codes based on new comments

* Fix bugs caused by shallow copy in attributes.py
* Imporve amend_distributed_attr_for_program in context.py
* Other changes for weihang's comments

* support shard reader

* support shard reader

* add parallel mode

* update process mesh

* add method to compute comm_group

* implement dist_embedding forward func

* implement dist matmul forward func

* implement dist reshape forward func

* add transpiler framework

* add transpiler forward

* implement transpiler forward

* implement transpiler backward & update

* add process

* add unitest

* chmod

* chmod

* chmod

* update unitest

* add unitest for gpt

* remove unused print

* rename transpiler --> partitioner

* rename transpiler --> partitioner

* chmod

* chmod

* bug fixed

* remove amp function

* update case for dp mode

* update case for dp mode

* [Auto Parallel] Integrate all parts with the newest code

* Integrate all parts of auto parallel and improve codes

* Integrate all parts by AutoParallelizer
* Add unit test for AutoParallelizer
* Improve auto completion module for pipeline parallel
* Add support for matmul_v2 in dist_matmul
* Correct the typo "stratergy" to "strategy"

* Modify distributed_strategy.proto to conform the main stream

* Restore parts of distributed_strategy to conform the develop branch
Co-authored-by: Nsandyhouse <lilong12@baidu.com>
Co-authored-by: NJZ-LIANG <jianzhongliang10@gmail.com>

12155358

refactor new executor (#35537) · 0eb7c942

由 wanghuancoder 提交于 9月 08, 2021

* refactor new executor, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop

0eb7c942

Enable program passes on Fleet APIs (#34955) · 5f369881

由 Zeng Jinle 提交于 9月 08, 2021

* add fleet api for program pass

* turn on apply pass for CI test

* fix disable fuse_all_optimizer bug

* try to test ci

* fix CI

* fill unspecified op role

* fix fuse_allreduce

* add ut to improve coverage

* remove useless change

* improve c++ coverage

* follow some comments

* test ir pass pipeline

* update doc

* reduce ut time again

5f369881

Work queue group (#35470) · a53460aa

由 liutiexing 提交于 9月 08, 2021

* Split Tracker and WorkQueue

* add WorkQueueGroup

* add unittest

* fix

* update

* update

* fix compile

a53460aa

W

[NPU] add get_float_status op and refine NPU check_nan_inf (#35274) · c727ec4a
由 WangXi 提交于 9月 08, 2021

c727ec4a

07 9月, 2021 2 次提交
- Y
  
  support multi-node (#35396) · c6e0cedc
  由 yaoxuefeng 提交于 9月 07, 2021
  
  c6e0cedc
- C
  
  fix int8 (#35504) · ed97be09
  由 ceci3 提交于 9月 07, 2021
  
  ed97be09
06 9月, 2021 1 次提交

Add fusion_lstm INT8 PTQ (#35334) · 7ef04da6

由 joanna.wozna.intel 提交于 9月 06, 2021

* Add fusion_lstm INT8 PTQ

* Correct mkldnn_cache_capacity and enable fc_lstm_fuse_pass only for this test

* Change mkldnn_cache_capacity

7ef04da6

03 9月, 2021 1 次提交

modify gc logic, use new device_event (#35208) · 80c0cc97

由 wanghuancoder 提交于 9月 03, 2021

* modify gc logic, use new device_event, test=develop

* use GenerateDeviceEventFlag, test=develop

* refine, test=develop

* fix test_standalone_executor.py, test=develop

* refine, test=develop

80c0cc97

01 9月, 2021 4 次提交
- T
  [HeterPs] merge dense && data norm && g2sum (#35029) · a647b80a
  由 Thunderbrook 提交于 9月 01, 2021
```
* merge dense

* log level

* tensor copy sync

* format
```
  a647b80a
- S
  [HybridParallel]Support finetinue model for PipelineParallel (#35287) · 264ff9ef
  由 ShenLiang 提交于 9月 01, 2021
```
* add cache for send_recv

* add eval_batch for pipeline

* add eval batch for pipelineparallel

* add style code
```
  264ff9ef
- W
  modify fetch logic, use D2H Stream (#35191) · c56d6978
  由 wanghuancoder 提交于 9月 01, 2021
```
* modify fetch logic, use D2H Stream, test=develop

* refine, test=develop
```
  c56d6978
- Q
  support KL label smooth (#35177) · 7ca28bb6
  由 QingshuChen 提交于 9月 01, 2021
```
* support KL label smooth

* update UT for KL label_smooth
```
  7ca28bb6
31 8月, 2021 2 次提交
- A
  Support CostInfo and MemProfiler in InterpreterCore (#34981) · 572bad8a
  由 Aurelius84 提交于 8月 31, 2021
```
* polish code

* fix unittest on windows

* refine pybind interface

* support statistic MemSize of AllocatorPool

* Replace mutex into atomic
```
  572bad8a
- 王
  
  fix the pass compat check position error, test=develop (#35272) · 54f07019
  由王明冬提交于 8月 31, 2021
  
  54f07019
30 8月, 2021 3 次提交
- C
  
  fix using boost::none as the init value when using paddle::optional (#35215) · e864667b
  由 chentianyu03 提交于 8月 30, 2021
  
  e864667b
- C
  [paddle-TRT]support matmul set to int8 in multihead (#34917) · 0043fa8c
  由 ceci3 提交于 8月 30, 2021
```
* update ernie int8
```
  0043fa8c
- A
  Abstract GenerateDeviceEventFlag to shield platforms (#35219) · 20cfa8ba
  由 Aurelius84 提交于 8月 30, 2021
```
* Abstract GenerateDeviceEventFlag to shield platforms

* Remove get_cuda_flags
```
  20cfa8ba
27 8月, 2021 2 次提交

Add fusion_gru and multi_gru to PTQ (Post-Training Quantization) (#33749) · 7debae3a

由 joanna.wozna.intel 提交于 8月 27, 2021

* Add calculation for gru op

* Correct the types

* Remove mkldnn only

* Correct mkldnn ifdef

* Remove mkldnn ifdef

* Separate mkldnn quantizer test

* Correct Windows test

* Check different cmake fix

* Revert cmake change

* Cmake change 2

* Cmake change 3

7debae3a

A
Polish DeviceEvent interface and Remove #ifdef in InterpreterCore (#35196) · 48bf7cbf
由 Aurelius84 提交于 8月 27, 2021
```
* add CPUDeiveEvent

* Polish DeviceEvent code

* Add DEVICE_EVENT_LIBS
```
48bf7cbf

26 8月, 2021 3 次提交

gc for newexecutor (#35085) · f1472039

由 wanghuancoder 提交于 8月 26, 2021

* gc for newexecutor, test=develop

* refine, test=develop

* add interpretercore_gc_helper.h,test=develop

* backup

* gc whit thread and device_event, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop

* fix bug, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop

* add CheckGC, test=develop

f1472039

Support Multi-Stream, Single-Thread in New Executor (#35024) · 678a259a

由 Aurelius84 提交于 8月 26, 2021

* Modify into QueueSync QueueAsync

* fix complie on MacOS

* fix pointer

* fix conflict

* polish unittest

* fix windows fetch error

* polish code according reviewer

* fix device_guard on CPU place

678a259a

W

[Inference] Replace unordered_map with map to support subgraph stability (#35147) · a1aae040
由 Wilber 提交于 8月 26, 2021

a1aae040

机器未来 / Paddle 与 Fork 源项目一致

机器未来 / Paddle
与 Fork 源项目一致