提交 · 93d862b0adf224a0af547d1442c57fbd6d0e8efc · 机器未来 / Paddle

24 8月, 2021 1 次提交

Add auto completion module for auto parallel (#34813) · 93d862b0

由 Yulong Ao 提交于 8月 24, 2021

* add auto_parallel dir

* mv to paddle.distributed

* add shard_xx api

* add distributed attrs for var

* add ut, test=develop

* add dist

* update

* update

* update

* update

* update

* update, test=develop

* update, test=develop

* update, test=develop

* update, test=develop

* update, test=develop

* update, test=develop

* update, test=develop

* update

* update

* update

* update

* update

* update, test=develop

* update, test=develop

* update

* update

* delete unused proto

* resotre op_desc

* restore type_defs

* update var_desc

* remove dimss_mapping for proto_pybind

* update interface.py

* update framework.py

* update

* update

* add auto_parallel dir

* mv to paddle.distributed

* add shard_xx api

* add distributed attrs for var

* add ut, test=develop

* [WIP] Add the auto completion feature and related codes

* [WIP] Improve the auto completion and related codes

* [WIP] Make the auto completion to support data-parallel

* [WIP] Make the completion support mp and dp+mp

* [WIP] Refactor auto completion unit test for MLP

* [WIP] Refactor the implementation of DistributedOperatorImpl

* [WIP] Improve dims_mapping update rule and fix a bug

* [WIP] Support auto completion for one transformer decoder layer

* [WIP] Add a minor change

* [WIP] Fix a bug within the uint test

* Shard XShape tensor, add embedding completion and refactor code

* Add the distributed_operators dir to setup.py.in

* Improve the completion process and add the unittest for gpt

* fix process_mesh ut

* fix process_mesh ut

* update

* update, test=develop

* Add support for automatically completing distributed attrs of special ops

* update

* update

* update

* fix doc sample codes, test=develop

* improve coverage, test=develop

* add static_mode check, test=develop

* Model the cluster for cost model and physical mapping

* update, test=develop

* add set_placement, test=develop

* Add the check to make sure the candidate tensors' size is great than zero

* update doc, test=develop

* update doc, test=develop

* update doc, test=develop

* update doc, test=develop

* update, test=develop

* Auto mark dist attrs annotated by user

* update ndarray to nested list, test=develop

* update, test=develop

* Add auto-completion module for auto-parallel (based on PR#33804)

* Remove unnecessary files

* Remove unrelated files for the auto completion pr

* Update the unit test to improve the coverage

* Modify codes based on reviews

* Minor changes for CI

* Improve some codes based on new comments

* Fix bugs caused by shallow copy in attributes.py
* Imporve amend_distributed_attr_for_program in context.py
* Other changes for weihang's comments
Co-authored-by: Nsandyhouse <lilong12@baidu.com>

93d862b0

23 8月, 2021 13 次提交
- B
  
  [CPU] Enable barrier op upon gloo (#34671) · e8f146a9
  由 Bo Liu 提交于 8月 23, 2021
  
  e8f146a9
- W
  
  trt convert ut add dynamic_shape and int8, etc. (#35061) · 17188e8d
  由 Wilber 提交于 8月 23, 2021
  
  17188e8d
- W
  
  remove old data check (#35077) · 5b814fd5
  由 wenbin 提交于 8月 23, 2021
  
  5b814fd5
- J
  [oneDNN] disable caching for interpolate and batch Norm (#35030) · 673bf719
  由 Jacek Czaja 提交于 8月 23, 2021
```
* - disabled interpolate onednn

* - compilation fix

* - draft of batch_norm cache disabling

* - fixes to UT
```
  673bf719
- P
  support infer_ut on windows nightly build (#35049) · 4f86aae0
  由 Peihan 提交于 8月 23, 2021
```
* enable infer_ut on windows

* remove lib calculation & time

* unset http_proxy when download bos file on windows
```
  4f86aae0
- L
  Refactor the organization of layer_norm cuda impl. (#34883) · 7f5eb533
  由 Li Min 提交于 8月 23, 2021
```
Refactor the organization of layer_norm cuda impl so that it can be reused in fused attention op.

    Extract the layer_norm cuda impl form layer_norm_op.cu to layer_norm_kernel.cu.h.
    Define fused/attention_layer_norm.h, which can be used in fused attention op in next PR.
```
  7f5eb533
- Z
  Support gettiem by Bool index (#35026) · b6dc16cb
  由 zyfncg 提交于 8月 23, 2021
```
* Support getitem by Bool index

* delete some debug info of bool index

* support the case that the shape of bool index is different from indexed tensor
```
  b6dc16cb
- W
  Revert "use spin lock in auto growth allocator (#34910)" (#35069) · 97fef015
  由 wanghuancoder 提交于 8月 23, 2021
```
This reverts commit 6bacfb0e.
```
  97fef015
- P
  
  add beam_search_decode npu op (#34967) · 4ce272ed
  由 pangyoki 提交于 8月 23, 2021
  
  4ce272ed
- P
  
  add fill_constant_batch_size_like npu op (#34563) · 7d86737c
  由 pangyoki 提交于 8月 23, 2021
  
  7d86737c
- T
  
  Fix a bug of strided_slice op, about the axes parameter access memory out of bounds (#35062) · aefec228
  由 TeslaZhao 提交于 8月 23, 2021
  
  aefec228
- S
  
  set node feature (#34994) · c3efabeb
  由 seemingwang 提交于 8月 23, 2021
  
  c3efabeb
- Z
  add adamw cuda kernel (#35020) · 77a8a394
  由 zhaoyingli 提交于 8月 23, 2021
```
* adamw support cuda

* adamw support cuda
```
  77a8a394
22 8月, 2021 1 次提交
- Z
  
  implementation of broadcast add backward by reduce (#34143) · 56c5e210
  由 Zhang Zheng 提交于 8月 22, 2021
  
  56c5e210
20 8月, 2021 10 次提交
- H
  
  Add paddle.linalg.matrix_power OP (#34667) · e2241a43
  由 Hao Lin 提交于 8月 20, 2021
  
  e2241a43
- Y
  
  [hybrid performance] Grad fuse for gradient merge under pipeline mode (#35004) · 4d9b2d6d
  由 Yuang Liu 提交于 8月 20, 2021
  
  4d9b2d6d
- L
  [npu]Add argsort op (#34865) · 99ffeffe
  由 lzzyzlbb 提交于 8月 20, 2021
```
* add rmsprop npu

* add argsort npu

* add argsort npu

* modify according to review

* modify sharedatawith according to review

* modify reshape according to review

* rm dygraph=false
```
  99ffeffe
- S
  [NPU] Support npu kernel for pad3d op (#34815) · ef517a56
  由 Sing_chan 提交于 8月 20, 2021
```
* [NPU] Support npu kernel for pad3d op

* fix for comment of zhouwei25

* fix some bugs according to qili93's comments

* add support and test for paddings in input

* delete VLOG used for debug
```
  ef517a56
- W
  use spin lock in auto growth allocator (#34910) · 6bacfb0e
  由 wanghuancoder 提交于 8月 20, 2021
```
* use spin lock in auto growth allocator, test=develop

* use pthread spin lock, test=develop

* use lock guard, test=develop

* use malloc spin lock, test=develop

* use lock_guard, test=develop
```
  6bacfb0e
- W
  fix set_lod in data_feed (#35000) · 4416c793
  由 wangguanqun 提交于 8月 20, 2021
```
* add trainer desc config to distributed strategy

* code style modified

* data_feed set lod
```
  4416c793
- Z
  [NPU] Support npu op depthwise_conv2d (#34853) · 4c115a82
  由 zhaoyingli 提交于 8月 20, 2021
```
* add depthwise_conv2d npu

* add some tests

* Delete test_unique_op_npu.py

* delete trans input
```
  4c115a82
- Z
  [NPU] Support npu op where and where grad (#34587) · d082955e
  由 zhaoyingli 提交于 8月 20, 2021
```
* [NPU] Support npu op where and where grad

* fix use const_cast

* delete a test
```
  d082955e
- P
  
  temporary disable resnet50-quant multi-thread test (#35035) · f927b653
  由 Peihan 提交于 8月 20, 2021
  
  f927b653
- J
  add (N,C,*) input support for GroupNorm (#34773) · 46371515
  由 JYChen 提交于 8月 20, 2021
```
* add (N,C,*) input support for GroupNorm

* --amend
```
  46371515
19 8月, 2021 6 次提交

[NPU] Support npu kernel for sin op (#34844) · 4641e8fc

由 JingZhuangzhuang 提交于 8月 19, 2021

* add npu sin op

* [NPU] Support npu kernel for sin op

* modify support npu kernel for sin op

* modify support npu kernel for sin op

* modify nou sin op

* modify npu sin op

* add sin op npu

4641e8fc

add resnet50_quant model in PR-CI-INFERENCE (#35012) · 97cae5e8

由 Peihan 提交于 8月 19, 2021

* add slim resnet50 quant model in pr-ci-inference

* enable resnet50_quant multi_thread4_trt_int8_bz1

* remove LOG(FATAL)

97cae5e8

Y
Add dimension check for inverse to avoid dividing by 0 error when input's... · a2e08657
由 Yiqun Liu 提交于 8月 19, 2021
```
Add dimension check for inverse to avoid dividing by 0 error when input's shape is [0, 0, 0]. (#34996)
```
a2e08657
C
fix batch_norm and instance norm when input is [] (#34107) · ca7f5208
由 ceci3 提交于 8月 19, 2021
```
* fix batch_norm and instance norm when input is []
```
ca7f5208

Fix Inference CI CPU/GPU (#34931) · 26213a77

由 tianshuo78520a 提交于 8月 19, 2021

* notest;test=gpu-inference

* notest;test=gpu-inference

* notest;test=gpu-inference

* notest;test=gpu-inference

* fix error

* notest;test=gpu-inference

* notest;test=gpu-inference

* notest;test=gpu-inference

* test=gpu-inference

26213a77

Abstract DeviceEvent to manage cross-platform Event implementation (#34922) · 22da1907

由 Aurelius84 提交于 8月 19, 2021

* add device_context

* add gtest for device_event_gpu

* Remvoe duplicate DeviceType

* push for test

* add unittest

* fix macros

* fix MSVC using usage

22da1907

18 8月, 2021 9 次提交
- L
  [NPU]add rmsprop op (#34864) · 9cbba97b
  由 lzzyzlbb 提交于 8月 18, 2021
```
* [npu]add rmsprop op
```
  9cbba97b
- X
  Add NPU kernel for norm Op: float16 and float32 (#34609) · 755c8a19
  由 xiongkun 提交于 8月 18, 2021
```
* Add NPU kernel for norm Op: float16 and float32

* fix code for code review

* fix for code review

* add type for paddle_throw

* remove unnecessary head file.\nAdd more testcase

* remove a broadcast
```
  755c8a19
- fix pad outliers err (#34979) · 248e27b7
  由 littletomatodonkey 提交于 8月 18, 2021
```
* fix pad outliers err

* fix pad api input type and doc

* fix example of pad

* add unittest for pad3d

* fix unittest

* fix error format

* fix pad doc
```
  248e27b7
- W
  code refactoring for new executor (#34970) · 40d4d834
  由 wanghuancoder 提交于 8月 18, 2021
```
* code refactoring, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop
```
  40d4d834
- P
  
  add paddle detection model in pr-ci-inference (#34986) · 1b747de7
  由 Peihan 提交于 8月 18, 2021
  
  1b747de7
- J
  [NPU] Add square grad (#34889) · 1b71a718
  由 Jackwaterveg 提交于 8月 18, 2021
```
* test=develop

* test=develop
```
  1b71a718
- J
  [NPU] Add leaky Relu (#34894) · 40f62737
  由 Jackwaterveg 提交于 8月 18, 2021
```
* test=develop

* test=develop
```
  40f62737
- W
  [Hybrid Performance] Move the cast op of AMP which cast fp32 param to fp16... · a9673b44
  由 WangXi 提交于 8月 18, 2021
```
[Hybrid Performance] Move the cast op of AMP which cast fp32 param to fp16 param to the optimizer (#34965)
```
  a9673b44
- C
  [CustomOp] Fix ext_tensor.cast failed bug (#34884) · 4d88cdb8
  由 Chen Weihang 提交于 8月 18, 2021
```
* fix ext_tensor.cast failed bug

* remove useless deps

* fix windows cmake failed

* try to fix windows make failed

* fix make error on windwos
```
  4d88cdb8

机器未来 / Paddle 与 Fork 源项目一致

机器未来 / Paddle
与 Fork 源项目一致