提交 · 5b357e021d33e02e16d5eb96d2f6d644c9c78277 · PaddlePaddle / Paddle

26 10月, 2021 11 次提交

[cherry-pick]Support FP16 in HybridParallel and Fix bugs in HybridOptimizer (#36707) · 5b357e02

由 Haohongxiang 提交于 10月 26, 2021

* fix bugs in HybridParallelClipGrad of hybrid_parallel_optimizer (#36237)

* fix bugs in HybridParallelClipGrad of hybrid_parallel_optimizer

* update

* update

* fix bugs in mp_layers、pp_layers and HybridParallelClipGrad (#36144)

* fix calling bug of HybridParallelClipGrad

* fix bugs of HybridParallelClipGrad

* add unittest of pp with HybridParallelClipGrad

* fix bugs in mp_layers.py

* update

* fix bugs in pp_layers.py

* update

* [HybridParallel]Rebuild code for pipeline (#36396)

* add no_sync for parameters sync

* add pipeline for moe

* [HybridParallel]Support fp16 in dygraph hybrid parallel (#36420)

* [HybridParallel]Support fp16 in dygraph hybrid parallel

* update

* update

* update for recompute

* add unittest of pp+fp16

* add unittest of recompute+fp16

* update

* modify ut

* modify ut of cond (#36475)

* fix bugs of ClipGradByGlobalNorm in HybridParallel (#36555)

* fix bugs of ClipGradByGlobalNorm

* add unittests

* add unittests

* [HybridParallel]fix bug of check_inf in fleet_base.py (#36651)

* fix bug of check_inf

* fix allreduce

* support ClipGradByGlobalNorm in sharding (#36012)

* support ClipGradByGlobalNorm in sharding

* support ClipGradByGlobalNorm in sharding

* test=allcase

* Update test_linalg_cond.py

* Update hybrid_parallel_util.py

* Update hybrid_parallel_util.py
Co-authored-by: NShenLiang <1422485404@qq.com>
Co-authored-by: Nzhaoyingli <86812880+zhaoyinglia@users.noreply.github.com>

5b357e02

Z
[cherry-pick]add op: fused_feedforward(forward) (#36729) · 77034fc3
由 zhangkaihuo 提交于 10月 26, 2021
```
This is a fusion operator to compute feed forward layer in transformer model architecture.
```
77034fc3
F

Pool3d 2.0 (#36545) (#36721) · dfda193f
由 feng_shuai 提交于 10月 26, 2021

dfda193f

Add bincount op (#36317) (#36709) · 610a810c

由 smallv0221 提交于 10月 26, 2021

* Add bincount op

* upload cpu version

* fix unitest

* fix unittest

* fix unittest

* fix en doc

* add more test

* fix en doc

* add more test case

* fix test

* fix input vailidation

* fix input check

* fix unittest

* fix test

* fix en doc

cherry-pick

610a810c

Y

[Cherry-pick] Add the forward QR operator (#36627) · 616ce203
由 Yulong Ao 提交于 10月 26, 2021

616ce203

Support various length support for SelectedRows in GLOO::AllGather (#36637) (#36722) · fced11bd

由 xiongkun 提交于 10月 26, 2021

Support various length support for SelectedRows in GLOO::AllGather (#36637)

    In cpu parallel using gloo, add various length support for SelectedRows

fced11bd

L
[Amp] refine code of amp level (#36362) (#36726) · 1ee4fc32
由 Leo Chen 提交于 10月 26, 2021
```
* refine amp level

* fix typo

* update tracer._amp_level
```
1ee4fc32
Y
add slot record support for GpuPS (#36723) · 53480c9c
由 yaoxuefeng 提交于 10月 26, 2021
```
* add slotrecord datafeed (#36099)

* fix multi-node (#36329)
```
53480c9c
H

cherry pick CrossEntropy's bug fix (#36647) · 32fe5a49
由 HydrogenSulfate 提交于 10月 26, 2021

32fe5a49

[cherry-pick-2.2] Fused attention op forward (#35905) (#36708) · d2be870a

由 Li Min 提交于 10月 26, 2021

功能：本PR的目标是提高attention模块的计算性能。
为了减少框架层对op的调度开销，本PR通过在C++层手动实现attention模块，对外提供attention 大op；
为了减少防存开销，本PR采取了两种优化方法：
（1）在q,k,v计算时通过共享输入X，将该处的gemm，transpose和bias add从三次调用减少为一次；
（2）使用kernel融合优化技术，在不同cuda kernel之间通过寄存器传输数据；

d2be870a

[cherry-pick] Support CPU Parallel in DataParallel Interface by GLOO to speed... · beb920cd

由 xiongkun 提交于 10月 26, 2021

[cherry-pick] Support CPU Parallel in DataParallel Interface by GLOO to speed up training (#35745) (#36605)

* User specified backend (#35745)

* remove tensordot

beb920cd

25 10月, 2021 9 次提交

[cherry-pick 2.2] static model parallel dropout support deterministic RandomSeedGenerator (#36682) · 59615fff

由 WangXi 提交于 10月 25, 2021

* Revert "Add fused_dropout wrapper to ease use. (#36185) (#36640)"

This reverts commit 05d7e2fd.

* [hybrid] seed and dropout op support force-cpu (#35820)

* [HIP] fix op not support AMD GPU bug, the flag PADDLE_WITH_ROCM is invalid

* [HIP] fix op not support AMD GPU bug, the flag PADDLE_WITH_ROCM is invalid

* [HIP] fix op not support AMD GPU bug

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] fix seed ci failed issue

* add AsExtra for force_cpu of seed op

* Add fused_dropout wrapper to ease use. (#36185)

* [hybrid] static model parallel dropout support deterministic RandomSeedGenerator (#36228)
Co-authored-by: Nxiayanming <41795079@qq.com>
Co-authored-by: NLi Min <11663212+limin2021@users.noreply.github.com>

59615fff

W

[cherry-pick] enable trt test check and fix trt ut error (#36549) (#36655) · a540769b
由 Wilber 提交于 10月 25, 2021

a540769b
W

[cherry-pick] enable trt test check and fix trt ut error (#36371) (#36654) · 0951bfd1
由 Wilber 提交于 10月 25, 2021

0951bfd1
J

Add pool2d test convert (#36338) (#36663) · 7612bf1c
由 JingZhuangzhuang 提交于 10月 24, 2021

7612bf1c
W

Cherrypick (#36666) · a9b7d1d2
由 wenbin 提交于 10月 25, 2021

a9b7d1d2

Add nn.functional.sparse_attention and some test cases, test=develop (#35757) (#36551) · c57d1e91

由 Liu-xiandong 提交于 10月 25, 2021

Add paddle.nn.functional.sparse_attention API

本个PR主要将sparse_attention功能在python层进行了一层封装，OP的主体代码见：#PR35676

此外，对于封装的python 接口，增加了相应的单测。

c57d1e91

Z
[Cherry Pick]Add fp16 kernel for clip_op (#36577) (#36672) · bd40dd9a
由 zhangbo9674 提交于 10月 25, 2021
```
Add fp16 kernel for clip_op.
```
bd40dd9a
Z
[Cherry Pick] refine comments for GradScaler state_dict (#36522) (#36671) · 304fb2b5
由 zhangbo9674 提交于 10月 25, 2021
```
Refine comments for GradScaler state_dict.
```
304fb2b5
F
[cherry-pick] Add new API 'tensordot' (#36273) (#36454) · 2bfee7d3
由 From00 提交于 10月 25, 2021
```
* Add new API tensordot
cherry-pick #36273
```
2bfee7d3

24 10月, 2021 1 次提交

Add viterbi decode (#35778) (#36615) · 1906c746

由 Jack Zhou 提交于 10月 24, 2021

* add viterbi decode cpu kernel

* add viterbi decoder api in paddle.text

* add a data buffer once to avoid create many small pieces of data buffer frequently

* fix viterbi max_seq_length bug

* fix seq_len=1 bug

* fix device context

* move split out of for loop

* remove INVERSE_SUB

* remove 2 GET_CAST_MASK

* remove 1 loop

* remove Functor

* add to_static deploy code

* use MAX_FUNC instead of ELE_MAX

* add MaxFunctor

* impl max_func

* remove MaxFunctor

* remove cast op

* use REGISTER_OP_WITHOUT_GRADIENT

* add viterbi cuda kernel

* add FIX_BLOCKDIM_CASE macro

* add MKL add, mul; add get data mask

* add arange mkl impl

* add CPU Argmax

* add cpu gather

* use EXECUTE_MKL_ELEMENT_BINARY_OP instead of some ADD, MUL

* use SameDimsBinaryOP instead of EXECUTE_MKL_ELEMENT_BINARY_OP

* use SAME_DIMS_ELEMENT_BINARY_OP

* add SimpleBroadcastBinaryOP

* use int instead of int64_t to accelerate

* optimize SimpleBroadcastBinaryOP

* optimize SimpleBroadcastBinaryOP

* optimize performance in both single thread and multithread situation

* remove useless line

* remove useless code

* add CREATE_TENSOR_BUFFER macro

* add INIT_REQUIRED_TENSOR macro

* add comment

* fix windows ci

* add viterbi unittest

* remove cuda add functor

* remove cuda equal

* remove a template function

* fix windows ci

* fix windows dtype

* remove some template instance

* remove useless header file

* remove some blockdim

* remove transpose impl

* accelerate cpu performance on single thread situation

* viterbi_decode->crf_decode

* rename crf params name

* add viterbi api test

* remove useless import

* add enable_static

* use viterbi decoder

* fix viterbi len=1

* fix  viterbi unittest

* remove useless comments

* reconstruct viterbi decode

* remove ADD,SUB,MUL structure

* fix coverage

* remove CREATE_TENSOR

* add name args

* crf.py->ops.py; with_start_stop_tag->include_start_end_tag

* update crf_decode en docs

* fix viterbi decode en docs

* fix some review comments

* add FIXED_BLOCK_DIM_CASE in cuda

* push_back->emplace_back

* crf_decode->viterbi_decode; include_start_end_tag->include_bos_eos_tag

* paddle.text.ops.viterbi_decode->paddle.text.viterbi_decode

* fix viterbi_decode en docs

1906c746

21 10月, 2021 2 次提交
- improve replicate pad error information (#36531) · a201a691
  由 littletomatodonkey 提交于 10月 21, 2021
```
* fix replicate pad when input size is 0

* add unit test
```
  a201a691
- 0
  remove no_value using var.name (#36513) (#36565) · 6a20205d
  由 0x45f 提交于 10月 21, 2021
```
* remove no_value using var.name
```
  6a20205d
20 10月, 2021 2 次提交
- W
  
  [cherry-pick] Inference add type check in copy_from_cpu (#36552) · b5404f09
  由 Wilber 提交于 10月 20, 2021
  
  b5404f09
- X
  catch the generatorfunction and intercept it. (#35369) (#36536) · 023eb3f9
  由 xiongkun 提交于 10月 20, 2021
```
* catch the generatorfunction and intercept it.

* add test generator

* add test case

* refine the testcase
```
  023eb3f9
19 10月, 2021 3 次提交

[cherry-pick]Add sparse attention cherrypick (#36447) · 36edb0e1

由 Liu-xiandong 提交于 10月 19, 2021

The code of this PR can only support CUDA 11.2. Currently, CI does not have GPU with CUDA 11.2 , and all tests will be skipped automatically.

The new OP is paddle._C_ops.sparse_attention. Regarding the work of the python API, it will be resolved in a follow-up PR.

The code of this PR lacks tests on dynamic graphs and static graphs, and will be added in subsequent PRs.

36edb0e1

C
quant support matmul_v2 (#36469) (#36499) · b8167ed2
由 ceci3 提交于 10月 19, 2021
```
* quant support matmul_v2

* fix format
```
b8167ed2

Add operators for async read & async write (#36333) (#36501) · d65f8af8

由 Siming Dai 提交于 10月 19, 2021

* fix async_read bug

* change index place to cpu

* add tensor size judge

* add async_read & async_write test

* fix bug in async_write

* fix mac py3 ci

* fix bug for cpu version paddle

* fix windows ci bug

* change input argument error type

* change const_cast to mutable_data

* add async_write out-of-bound check and consumate error hint

* fix a small bug for dst_tensor

* add docs and refine codes

* refine docs

* notest,test=windows_ci

* fix windows ci

* fix require

* fix code-block

* add core.is_compiled_with_cuda()

d65f8af8

18 10月, 2021 1 次提交

[Cherry-pick][Dy2stat]fix no_grad context error in train mode when using... · 2b9d1922

由 0x45f 提交于 10月 18, 2021

[Cherry-pick][Dy2stat]fix no_grad context error in train mode when using save/load (#36434) (#36463)

修复使用jit.save/load接口加载模型后，在train模式和no_grad上下文中，显存会一直增长的问题

2b9d1922

15 10月, 2021 2 次提交

[cherry-pick]Verify the correctness of graph rewrited by GeneratePass (#36453) · cc449652

由 wuhuanzhou 提交于 10月 15, 2021

* [WIP]Verify the correctness of graph rewrited by GeneratePass, test=develop

* add delete subgraph and unittest, test=develop

* check simple pass, test=develop

* fix coverage, test=develop

* limit with input_spec via Paddle API, test=develop

cc449652

Y
[cherry-pick] add sparse_embedding doc (#36312) · fc429fea
由 Yanxing Shi 提交于 10月 15, 2021
```
* add sparse_embedding doc

* modify sample code

* fix sample code error
```
fc429fea

14 10月, 2021 1 次提交
- fix windows bug that python virtual env can't find python executable (#36227) (#36370) · 976f0146
  由 zhouweiwei2014 提交于 10月 14, 2021
```
ATT，cherry-pick #36227
```
  976f0146
13 10月, 2021 3 次提交
- 0
  delete remove_static_file() function in error.py (#36153) (#36375) · a5767bb6
  由 0x45f 提交于 10月 13, 2021
```
* change time to remove static tempfile

* delete remove_static_file() function
```
  a5767bb6
- W
  [cherrypick] change paddle.mm api to matmul v2 op (#36374) · 7a66160d
  由 wawltor 提交于 10月 13, 2021
```
* change the paddle.mm to matmul_v2

* update the code for the mm

* update the document for the mm
```
  7a66160d
- J
  
  fix for matmul_v2 6D x 2D (#36379) · ce6a27d9
  由 jakpiase 提交于 10月 13, 2021
  
  ce6a27d9
12 10月, 2021 1 次提交
- A
  Fix stop_gradient in RunProgramOp (#36339) (#36353) · a6868c91
  由 Aurelius84 提交于 10月 12, 2021
```
* Fix stop_gradient in RunProgramOp

* fix reference
```
  a6868c91
11 10月, 2021 2 次提交
- S
  
  dlpack fix (#35817) (#36177) · 31a5829a
  由 Siming Dai 提交于 10月 11, 2021
  
  31a5829a
- W
  [cherry-pick]fix hasattr(paddle.fluid.ir.PassDesc.OP, '__name__') error (#36294) · 45de9312
  由 wuhuanzhou 提交于 10月 11, 2021
```
对于__getattr__重载后不满足条件的参数，全部抛出AttributeError异常，达到与未重载版本一致。

(cherry picked from PR #36229)
```
  45de9312
30 9月, 2021 2 次提交

Z
add optest for adamw (#36148) (#36239) · 70e67843
由 zhaoyingli 提交于 9月 30, 2021
```
* update func name

* skip cpu

* update unittest

* update unittest
```
70e67843

李

Fix raw optim (#36176) (#36231) · 28d12007

由李季提交于 9月 30, 2021

* fix raw optim

* pre-commit test file
Co-authored-by: Nsneaxiy <sneaxiy@126.com>
Co-authored-by: Nsneaxiy <sneaxiy@126.com>

28d12007

PaddlePaddle / Paddle 1 年多 前同步成功

PaddlePaddle / Paddle
1 年多前同步成功