提交 · 1b83de2eff362117507794ffaab6f6c29ba07809 · PaddlePaddle / Paddle

23 4月, 2021 6 次提交
- Z
  
  update 2.0 public api in optimizer (#31944) · 1b83de2e
  由 zhiboniu 提交于 4月 23, 2021
  
  1b83de2e
- B
  solve hccl communicate conflict (#32447) · 0e74eea2
  由 Baibaifan 提交于 4月 23, 2021
```
solve hccl communicate conflict (#32447)
```
  0e74eea2
- L
  add c_concat and c_split ops (#32486) · 2b108a04
  由 lilong12 提交于 4月 23, 2021
```
* add c_concat op
```
  2b108a04
- S
  
  add lstm support on xpu test=kunlun (#32436) · b6f8ccd2
  由 shanliang1992 提交于 4月 23, 2021
  
  b6f8ccd2
- S
  
  disable utest (#32474) · 1dc83932
  由 ShenLiang 提交于 4月 23, 2021
  
  1dc83932
- K
  Fix seven error message (#32397) · 203ac4f3
  由 Kqnonrime 提交于 4月 23, 2021
```
* fix two error message

* fix two error message

* fix error

* fix error

* fix error

* fix error

* fix some error message

* fix some error

* fix error

* fix some error

* fix some error

* fix some error

* fix one error

* fix some error

* fix seven error message

* fix error

* fix error

* fix error

* fix error
```
  203ac4f3
22 4月, 2021 12 次提交
- Y
  
  Add `paddle.set_grad_enabled` (#31794) · f8ca5a9d
  由 Yang Zhang 提交于 4月 22, 2021
  
  f8ca5a9d
- W
  support int32 and int64 kernel for clip operator (#32373) · c3328288
  由 wuyefeilin 提交于 4月 22, 2021
```
support int32 and int64 kernel for clip operator 
```
  c3328288
- H
  
  fix doc for adamw (#32438) · c4815707
  由 hutuxian 提交于 4月 22, 2021
  
  c4815707
- Y
  
  Add fleet get_loss_scaling doc and update alert message (#32419) · d03b0b16
  由 Yuang Liu 提交于 4月 22, 2021
  
  d03b0b16
- F
  import sequence_* API to new namespace (#32089) · f12c943a
  由 Feiyu Chan 提交于 4月 22, 2021
```
* import sequence_* API to new namespace

* fix typos, remove alias marking

* update sample code

* fix sample code

* fix docstring for sequence_mask
```
  f12c943a
- W
  modify conv2d_transpose docs (#32410) · 1064f2b8
  由 wangxinxin08 提交于 4月 22, 2021
```
* modify conv2d_transpose docs
```
  1064f2b8
- Z
  
  fix type(x)=paddle.VarBase to paddle.Tensor (#32364) · bec4b167
  由 zhiboniu 提交于 4月 22, 2021
  
  bec4b167
- S
  [HybridParallel] Add ClipGradByGlobalNorm & check_finite_and_unscale in Dygraph (#32354) · 7ea999fd
  由 ShenLiang 提交于 4月 22, 2021
```
* add clip/check

* add amp & clip grad in dygraph

* add logging
```
  7ea999fd
- F
  add glu in nn.functional (#32096) · b2ee8380
  由 Feiyu Chan 提交于 4月 22, 2021
```
add glu in nn.functional
```
  b2ee8380
- W
  
  strip after compilation (#32145) · e727820d
  由 wuhuanzhou 提交于 4月 22, 2021
  
  e727820d
- W
  support save/load binary format tensor. (#32211) · f4d9adc7
  由 WeiXin 提交于 4月 22, 2021
```
* support save/load binary format tensor

* Fix error when create cudaplace

* Fix error when create cudaplace

* Fix error when create cudaplace

* get devive context from pool.

* move define of 'SerializeToStream' and 'DeserializeFromStream' to 'lod_tensor.cc' and 'selected_rows.cc'.

* improve coverage.

* improve coverage.

* polish API

* deal with conflict

* disable save/load large file in unnittest

* split unnittest.
```
  f4d9adc7
- T
  
  Delete WITH_GRPC flag and Distributed old code (#32383) · e58c705b
  由 tianshuo78520a 提交于 4月 22, 2021
  
  e58c705b
21 4月, 2021 13 次提交

C
[HotFix] Add support for optimizer with varbase input (#32362) · b47dd158
由 Chen Weihang 提交于 4月 21, 2021
```
* add support for optimizer with varbase input

* refine cond

* fix failed unittest

* add test for coverage
```
b47dd158

【NPU】Merge NPU ccl code (#32381) · c3158527

由 zhang wenhui 提交于 4月 21, 2021

* add allreduce and broadcast without test (#31024)

add allreduce and broadcast without test

* Refactor HCCLCommContext to be compatible with Paddle (#31359)

Refactor HCCLCommContext to be compatible with Paddle (#31359)

* [NPU] add npu kernel for communication op (#31437)

* add allreduce and broadcast without test

* add c_broadcast_test case

* build c_comm_init and c_create_group operators

* make the whole thing compile

* add broadcast and init op test case but run failed

* make unit test compile

* fix broadcast test bug and change into hcom for ccl

* change c_comm_init and c_create_group ops accordingly

* make tests compile

* transfer code to 27

* compiled successfully in 28, but run failed

* test broadcast in 28, but failed

* make hcom primitives work

* change hccl data type for base.h

* fix broadcast bug

* make attributes work

* fix group name bug

* add allreduce but test failed

* allreduce bug for qiuliang

* allreduce finished

* add allgather and reducescatter

* merge all op code

* add allgather test

* finish run all ccl op test exclude send/recv

* all all op and test exclude send/recv

* send_v2_npu.cc recv_v2_npiu.cc compiled

* fix ccl core dump bug and test allgather, reducescatter, broadcast op

* fix allreduce bug just for test

* hcom send&recv test pass, without hcom_destroy

* for qiuliang test

* Ascend Send&Recv Test Pass

* all op (ex send/recv) ok

* fix bug

* merge all ccl op

* style merge to PaddlePaddle

* merge style

* new merge style

* merge style 2

* insert an empty at the end

* disable ctest for hcom to pass ci
Co-authored-by: Nvoid-main <voidmain1313113@gmail.com>
Co-authored-by: Nf2hkop <f2huestc@outlook.com>

* Add auto-increasing tag id for Hcom OPs (#31702)

* add c_reduce_sum op (#31793)

add c_reduce_sum op

* update Ascendrc hccl to 20.3 (#32126)

update Ascendrc hccl to 20.3 (#32126)

* fix merge code

* change cmake.txt1

* [NPU] Support npu kernel for c sync stream op (#31386)

* sync stream npu op

* add with_ascend_acl

* update c++ unittest

* compile all failed

* try to pre commit

* after pre commit

* merge&compile&test hccl successfully!

* fix code style

* fix code style

* fix bugs about hccl

* fix some bugs

* fix code style

* fix style

* fix style

* fix

* fixed

* merge develop
Co-authored-by: Nlw921014 <liuwei921014@yeah.net>
Co-authored-by: NVoid Main <voidmain1313113@gmail.com>
Co-authored-by: Nf2hkop <f2huestc@outlook.com>
Co-authored-by: Nxiayanming <41795079@qq.com>

c3158527

Y

Do not define and save reserve_space for inference. (#32375) · bc90916e
由 Yiqun Liu 提交于 4月 21, 2021

bc90916e
H

fix bug in amp O2 (#32343) · 4be3b057
由 huangxu96 提交于 4月 21, 2021

4be3b057
A

[CustomOp]Fix MAC3-CI random failed with XXX_setup.py(#32369) · 7bae5e9a
由 Aurelius84 提交于 4月 21, 2021

7bae5e9a
A

[CustomOP]Support find include/c++/v1 include dirs automatically (#32404) · 661a1f6f
由 Aurelius84 提交于 4月 21, 2021

661a1f6f
Y

add get_loss_scaling to fleet (#32401) · 37bb3342
由 Yuang Liu 提交于 4月 21, 2021

37bb3342
L
[NPU] register npu finalize on exit (#32390) · 8e4c1936
由 Leo Chen 提交于 4月 21, 2021
```
* [NPU] register finalize on exit

* fix
```
8e4c1936
L

[Kunlun]add collective ops for multi XPU cards training and add Kunlun multi XPU cards CI (#32302) · 2194ad15
由 liuyuhui 提交于 4月 21, 2021

2194ad15
J

Added bilinear and nearest interp v2 oneDNN FP32 kernels (#32312) · 5d19f8d8
由 jakpiase 提交于 4月 21, 2021

5d19f8d8
G

add test=develop (#32380) · 4898c38d
由 gongweibao 提交于 4月 21, 2021

4898c38d
J

Added oneDNN reduce_op GRAD kernel (#32280) · ead83422
由 jakpiase 提交于 4月 21, 2021

ead83422
X
remove fluid for auto_checkpoint. (#32157) · 1593ee25
由 xiemoyuan 提交于 4月 21, 2021
```
* remove fluid for auto_checkpoint.

* fix bug.
```
1593ee25

20 4月, 2021 4 次提交
- F
  add paddle.nn.unfold #32297 (#32298) · 186682fe
  由 FNRE 提交于 4月 20, 2021
```
* add paddle.nn.unfold
* update Parameters of Unfold
```
  186682fe
- J
  [Sharding]: update config DOC (#32299) · e3489013
  由 JZ-LIANG 提交于 4月 20, 2021
```
* sharding: update config DOC

* update pipeline config

* sharding update doc
```
  e3489013
- W
  
  save/load program (#32336) · e0a52fd7
  由 WeiXin 提交于 4月 20, 2021
  
  e0a52fd7
- W
  
  support `numpy.array/asarray(tensor) -> ndarray`, test=develop (#32300) · 43926c80
  由 Wenyu 提交于 4月 20, 2021
  
  43926c80
19 4月, 2021 4 次提交

[NPU] cherry-pick gc/dataloader/save&load/optimization from ascendrc to develop (#32294) · cbe5c9f8

由 Leo Chen 提交于 4月 19, 2021

* [NPU] support GarbageCollector for npu (#31874)

* support GarbageCollector for npu

* fix typo

* fix gather_grad

* disable NPUDefaultStreamGarbageCollector on NPU

* [NPU] support npu for memcpy op (#31808)

* support npu for memcpy op

* add ut

* fix ut

* fix typo

* 【NPU】fix bug of using temp vector (#31963)

* fix bug when beta1_pow on cpu (#31995)

* [NPU] support npu profiler (#31684)

* support npu profiler

* add python api

* fix bugs

* add wrapper for incomplete type

* update profile proto

* record npu wait

* add xpu placeholder

* fix adam (#32016)

* [NPU] enable async copy and  add wait before sync operation (#31956)

* enable async copy and  add wait before sync operation

* remove unneccessary wait

* add FillNpuTensorWithConstant

* refine

* fix fill_constant

* make TensorFromVector/TensorToVector sync

* [NPU] Support dataloader on npu place. (#31867)

* [NPU] Wait on NPUPlace (#32086)

* [NPU] fix cast op (#32121)

* fix npu kernel of cast op to handle casting to same dtype

* add comments

* [NPU] support cann 20.3 (#32044)

* fix compile problem on cann 20.3

* fix ut

* fix test_mul

* fix check_finite_and_scale

* fix lookup_table_v2_grad

* fix cmake

* support print op

* [NPU] Support npu save load (#31893)

* support save load for NPU

* add save load npu unittest

* support np.array transform in NPU

* fix errors

* delete dygraph in unittest

* add Wait

* fix unittest

* fix review comment

* fix unittest problem

* fix little problem

* change aclrtSynchronizeDevice to aclrtSynchronizeStream for better performance (#32196)

* change aclrtSynchronizeDevice to aclrtSynchronizeStream for better performace

* refine code

* fix NPUDeviceContext in all c++ unittest (#32198)

* fix NPUDeviceContext in all c++ unittest

* refine log
Co-authored-by: Npangyoki <pangyoki@126.com>

* [NPU] Remove TensorFromVector and avoid sync copy in npu op kernel for better performance (#31994)

* enable async copy and  add wait before sync operation

* remove unneccessary wait

* add FillNpuTensorWithConstant

* refine

* fix fill_constant

* change TensorFromVector to FillNpuTensorWithConstant

* fix ignored api

* delete extra unittest

* fix little error

* fix update_loss_scaling_op_npu and check_finite_and_unscale_op_npu

* change TensorCopySync to TensorCopy

* delete useless Wait and add StreamWait

* fix npu_stream error

* fix check_finite_and_unscale_op_npu TensorCopy

* only save stream wait

* fix NPUDeviceContext in all c++ unittest

* delete wait
Co-authored-by: Nzhiqiu <chenqiuliang@baidu.com>

* delete useless unittest file (#32206)

* Fix op test (#32231)

* fix conditional block (#32243)

* fix adam bug again (#32246)

* fix compile

* fix ut

* fix ut
Co-authored-by: Nliym27 <33742067+liym27@users.noreply.github.com>
Co-authored-by: Npangyoki <pangyoki@126.com>

cbe5c9f8

S
[Hybrid Parallel] Support dp & mp in dygraph (#32323) · ffd40860
由 ShenLiang 提交于 4月 19, 2021
```
* support dp & mp
```
ffd40860

Fix sublayer (#31824) · 4d69eeaa

由 Jiabin Yang 提交于 4月 19, 2021

* fix sublayer error with include_sublayers=False

* add ut

* refactor include_sublayers related api

* fix ut

* fix ut of transformer

* fix ut of transformer

* remove useless code

* change sublayer api

* polish code

* add test for include_self=True

4d69eeaa

J

Add BF16 Constant Initializer and support for other initializer (#31935) · 76cb83e8
由 joanna.wozna.intel 提交于 4月 19, 2021

76cb83e8

17 4月, 2021 1 次提交
- S
  [Hybrid Parallel] Add model parallel support in dygraph (#32248) · 66d46221
  由 ShenLiang 提交于 4月 17, 2021
```
* add model parallel support in dygraph
```
  66d46221

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功