提交 · revert-52175-dev_peak_memory · PaddlePaddle / Paddle

18 4月, 2023 1 次提交
- L
  Revert "fix peak memory (#52175)" · e227d093
  由 LiYuRio 提交于 4月 18, 2023
```
This reverts commit bd3b6adf.
```
  e227d093
14 4月, 2023 1 次提交
- J
  Eb118 BF16 Adoption (#52827) · 6f3c9643
  由 JZ-LIANG 提交于 4月 14, 2023
```
* pr1

* pr2

* pr3

* fixed unitest

* adopt for scale
```
  6f3c9643
12 4月, 2023 3 次提交
- Y
  
  Cherry-pick the support of bf16 of grad_clip, in #51285. (#52816) · 8cbc75ca
  由 Yiqun Liu 提交于 4月 12, 2023
  
  8cbc75ca
- J
  
  Cherry Pick Random Ctrl (#52778) · 3869a3b4
  由 JZ-LIANG 提交于 4月 12, 2023
  
  3869a3b4
- Y
  
  Unify the static amp codes of fp16 and bf16. Reimplement #52694 in release/2.4. (#52697) · 6959eae5
  由 Yiqun Liu 提交于 4月 12, 2023
  
  6959eae5
11 4月, 2023 1 次提交

Cherry pick for fix of operator precision. (#52705) · d1e8b1e2

由 Yiqun Liu 提交于 4月 11, 2023

* Fix scale kernel for low precision, cherry pick #50998.

* Fix the FP16 precision problem of add_n. (#50129)

* Change squared_l2_norm to reuse ReduceKernel, and register fp16 and bf16 kernel, which is cherry pick #48315.

* Cherry-pick the fix of MPTypeTrait in KP, which is implemented in #50993.

* Cherry-pick the multi-precision support of AdamW for bf16, #48041.

* Fix compiling error.

* Cherry-pick the fix of CubTensorReduceImpl for bfloat16 in #50993.

* Fix unittest.

---------
Co-authored-by: Nliuruyan <44316842+liuruyan@users.noreply.github.com>

d1e8b1e2

10 4月, 2023 1 次提交
- Y
  Broadcast the master weight along with param for distributed training. (#52638) · d12588d2
  由 Yiqun Liu 提交于 4月 10, 2023
```
* Broadcast the master weight along with param for distributed training.

* Fix codestyle.
```
  d12588d2
09 4月, 2023 3 次提交

Add bfloat16 support for several operators and apis. (#52696) · ba9a22db

由 Yiqun Liu 提交于 4月 09, 2023

* Cherry-pick the register of bfloat16 for amp_kernel, pull request #45541.

* Cherry-pick the master_grad support of adamw, pull request #51141.

* add bf16 for some ops in static mode (#51582)

* Add bfloat16 support for some api in static mode.

* Fix codestyle.

* Revert the change of layer_function_generator.py.

---------
Co-authored-by: Shaojie WANG <wsjmessi@163.com>

ba9a22db

Cherry pick the support of bfloat16 for several operators. (#52608) · 95c3d613

由 Yiqun Liu 提交于 4月 09, 2023

* Register exp/expm1/logit bf16 activation op kernels (#48702)

* register more bf16 ops

* update to register coresponding backward ops

* Addition of bf16 type support for Compare OP  (#46413)

* first commit

* clarify the quotes

* change code style format

* support bfloat16

* add bfloat16 support for more ops (#48272)

* [Bfloat16]register bfloat16 datatype for squared l2 norm (#50908)

* Sync the pull request #51903.

* Add some header files back.

* modify cmake file for cuda11.8 compile (#49020)

* modify cmake file for cuda11.8 compile

* add op_library(fused_embedding_eltwise_layernorm_op DEPS bert_encoder_functor)

* Fix compling error.

* Cherry-pick pull request #51396.

---------
Co-authored-by: Nsneaxiy <32832641+sneaxiy@users.noreply.github.com>
Co-authored-by: Nlimingshu <61349199+JamesLim-sy@users.noreply.github.com>
Co-authored-by: Shaojie WANG <wsjmessi@163.com>
Co-authored-by: Nzqw_1997 <118182234+zhengqiwen1997@users.noreply.github.com>

95c3d613

07 4月, 2023 1 次提交
- Z
  
  modify cmake file for cuda11.8 compile (#49020) (#52481) · 9431bae1
  由 zhaoyingli 提交于 4月 07, 2023
  
  9431bae1
03 4月, 2023 1 次提交
- Z
  
  make micro bsz configurable (#52447) · 722f880e
  由 zhaoyingli 提交于 4月 03, 2023
  
  722f880e
30 3月, 2023 1 次提交
- Y
  
  use int64 for c split (#52279) · 964497b5
  由 Yuang Liu 提交于 3月 30, 2023
  
  964497b5
28 3月, 2023 1 次提交
- L
  
  fix peak memory (#52175) · bd3b6adf
  由 LiYuRio 提交于 3月 28, 2023
  
  bd3b6adf
24 3月, 2023 1 次提交
- L
  
  optimize overlap between steps (#51974) · 81f4ef4f
  由 LiYuRio 提交于 3月 24, 2023
  
  81f4ef4f
20 3月, 2023 1 次提交
- L
  
  Cherry-pick fleet executor and auto parallel (#50071) · 92c2dcbd
  由 LiYuRio 提交于 3月 20, 2023
  
  92c2dcbd
09 3月, 2023 1 次提交
- J
  
  Extra Sync for Tensor Parallel (#50637) · 4bacf2ab
  由 JZ-LIANG 提交于 3月 09, 2023
  
  4bacf2ab
17 2月, 2023 1 次提交
- W
  
  Add rpc ops to fetch data from remote service (#50220) · 9025fddd
  由 Wen Sun 提交于 2月 17, 2023
  
  9025fddd
13 1月, 2023 2 次提交
- X
  
  fix_arg_release24 (#49771) · 0699afb1
  由 xiaoxiaohehe001 提交于 1月 13, 2023
  
  0699afb1
- Y
  fix fc kernel diff (#49781) · 01c26ab2
  由 Yuanle Liu 提交于 1月 13, 2023
```
* fix fc kernel diff

* disable fc_elementwise_layernorm_fuse_pass
```
  01c26ab2
12 1月, 2023 1 次提交
- X
  
  fix_split_infermeta (#49745) · 8a934047
  由 xiaoxiaohehe001 提交于 1月 12, 2023
  
  8a934047
09 1月, 2023 1 次提交
- H
  
  fix bugs of paddle.multiplex API (#49368) (#49642) · 6d2d8e50
  由 Haohongxiang 提交于 1月 09, 2023
  
  6d2d8e50
04 1月, 2023 2 次提交
- Y
  [Cherry-pick][Paddle Inference] fix mixed precision diff (#49477) · 1d25c663
  由 Yuanle Liu 提交于 1月 04, 2023
```
* disable scale op in amp pass

* Do not insert redundant cast op

* fix fused_fc_elementwise_layernorm kernel diff

* fix fc kerenl diff
```
  1d25c663
- Y
  [Cherry-pick] add condition of skipif (#49407) · 7696ae02
  由 YUNSHEN XIE 提交于 1月 04, 2023
```
* resolve conflict

* fix format error
```
  7696ae02
03 1月, 2023 2 次提交
- X
  [Cherry pick] fix fold for big bs (#49491) · 2a438b0a
  由 xiaoting 提交于 1月 03, 2023
```
* fix fold for large bs

* fix fold for large bs

* fix pre-commit
```
  2a438b0a
- F
  cherry-pick:Some version of TensorRT don't support qkv_plugin (#49425) · d7855fe8
  由 feng_shuai 提交于 1月 03, 2023
```
* cherry-pick:Some version of TensorRT don't support qkv_plugin

* cherry-pick:support coverage CI
```
  d7855fe8
30 12月, 2022 1 次提交

[MLU] cherry-pick from develop to release/2.4 (#48313) · 6e154fc6

由 Chenxiao Niu 提交于 12月 30, 2022

* [MLU] fix compute error of dropout op (#45923)

* [MLU] add mergedAdam kernel. (#45965)

* [MLU] add int64 support for mlu one_hot_v2 (#46313)

* [MLU] fix profiler compile failure (#46208)

* [MLU] add barrier_op kernel. (#46417)

* [MLU] fluid: add mluop (#46429)

* [MLU] add huber_loss kernel. (#46455)

* [MLU] add mlu kernel for add_reduce_max_grad (#45651)
Co-authored-by: Nliupeiyu <liupeiyu@cambricon.com>

* [MLU] add_fluid_mluop_yolo_box (#46573)

* [MLU] fix phi::Tensor compile error of mlu. (#46649)

* [MLU] add fluid MLUOps prior_box (#46585)

* [MLU] fix cmake error (#46772)

* [MLU]fix unittest of sync_bn (#46797)

* [MLU] add masterparam support for mlu adamw. (#46804)

* [MLU] add int64 support for allgather. (#46830)

* [MLU] fix compile error & add mlu blacklist function. (#47439)

* [MLU] fix softmax_with_cross_entropy failed in 370-X8.

* [MLU] fix cncl stuck caused by multiple initializations.

* [MLU] fix code style check.
Co-authored-by: Nqipengh <huangqipeng@cambricon.com>
Co-authored-by: Ncifar10 <41565156+cifar10@users.noreply.github.com>
Co-authored-by: Lux et Veritas <1004239791@qq.com>
Co-authored-by: Nliupeiyu <liupeiyu@cambricon.com>
Co-authored-by: Nronnywang <ronny1996@163.com>

6e154fc6

29 12月, 2022 2 次提交
- [cherry-pick]fix bug of UT test_version, test=document_fix (#49401) · 96e974a0
  由 zhouweiwei2014 提交于 12月 29, 2022
  
  96e974a0
- Y
  [Cherry-pick]Move sum op to PHI && Fix MetaTensor's bug when run infermeta (#49342) · 8015fbd6
  由 YuanRisheng 提交于 12月 29, 2022
```
* cherry-pick 45860

* [BUG FIX]Fix MetaTensor's bug when run infermeta (#46265)

* fix sum bug

* fix ci bugs

* fix ci bugs

* update code according comment
```
  8015fbd6
28 12月, 2022 1 次提交
- H
  [Cherry-pick] Fix CUDA11.8 Unittest Accuracy (#49374) · 8aa5be90
  由 Huihuang Zheng 提交于 12月 28, 2022
```
Fix CUDA11.8 Unittest Accuracy
```
  8aa5be90
27 12月, 2022 2 次提交

Y

update jetson ampere sm (#49364) · b5fdd175
由 Yuanle Liu 提交于 12月 27, 2022

b5fdd175

[Cherry-pick] Fix custom operator backward=None (#48656) (#48715) · 39eb77a6

由 HongyuJia 提交于 12月 27, 2022

* [Release2.4] Revert python link prs (#48573)

* Revert "Fix mac link python (#48017)"

This reverts commit 3fa7a736.

* Revert "[Cherry-pick] Fix python link error (#47811)"

This reverts commit ff642c68.

* Update config.go

* fix custom operator backward=None (#48656)

* [Custom Extension] Fix custom double_grad backward=None (#49224)

* fix custom double_grad backward=None

* fix custom_relu.cu bug && polish testcase of double_grad

* remove old dynamic graph test

* add import fluid

* add import fluid
Co-authored-by: NChen Weihang <chenweihang@baidu.com>

39eb77a6

22 12月, 2022 3 次提交

G

fix unittest in post training quantization (#49257) · 5d29a5bf
由 Guanghua Yu 提交于 12月 22, 2022

5d29a5bf

Fix mixed precision bug (#49239) · 11c7f570

由 Yuanle Liu 提交于 12月 22, 2022

* [Release2.4] Revert python link prs (#48573)

* Revert "Fix mac link python (#48017)"

This reverts commit 3fa7a736.

* Revert "[Cherry-pick] Fix python link error (#47811)"

This reverts commit ff642c68.

* Update config.go

* fix mixed precision inference
Co-authored-by: NChen Weihang <chenweihang@baidu.com>

11c7f570

L

[Docs]update readme; test=document_fix (#49246) · 612bdb17
由 Ligoml 提交于 12月 22, 2022

612bdb17

21 12月, 2022 2 次提交
- A
  
  fix unittests (#49203) (#49210) · 7c36b887
  由 Aganlengzi 提交于 12月 21, 2022
  
  7c36b887
- Z
  
  cherry-pick #75b734 (#49201) · fb19648a
  由 zhangkaihuo 提交于 12月 21, 2022
  
  fb19648a
20 12月, 2022 1 次提交
- S
  Fix nullptr to TestFuseGemmEpilogueReluBWDFP* (#48997) (#49090) · cdab3a44
  由 ShenLiang 提交于 12月 20, 2022
```
Co-authored-by: NMing-Xu Huang <mingh@nvidia.com>
```
  cdab3a44
19 12月, 2022 1 次提交

[cherry-pick][Inference] support mixed precision inference (#49077) · ddcd1b61

由 Yuanle Liu 提交于 12月 19, 2022

* [Release2.4] Revert python link prs (#48573)

* Revert "Fix mac link python (#48017)"

This reverts commit 3fa7a736.

* Revert "[Cherry-pick] Fix python link error (#47811)"

This reverts commit ff642c68.

* Update config.go

* [Paddle Inference] Add float_to_half_pass to support  inference with mixed precision (#47993)

* [Inference] optimize some code and fix some bug (#48780)

* clean ir_pass_manager and fix map_depthwise_conv_to_conv_pass

* fix unitest timeout

* [Paddle Inference] clean unused code  (#48392)

* fix

* update

* update
Co-authored-by: NChen Weihang <chenweihang@baidu.com>

ddcd1b61

29 11月, 2022 1 次提交

[cherry-pick] updating mul and matmul with set_mem_desc and fix... · 9e2ba9b9

由 yeliang2258 提交于 11月 29, 2022

[cherry-pick] updating mul and matmul with set_mem_desc and fix squeeze_transpose for MKLDNN (#47951)

* Fix slice bugs in MKLDNN when input dims are zeros (#46671)

* fix slice bugs

* fix

* update code

* fix

* update code

* updating mul and matmul with set_mem_desc (#45624)

* - mul & matmul changes

- fix

- bs16 correction of strides

* - cosmetic fixes

* - lint

* - fix

* - fix

* - format -> mem_desc

* - fix

* - fix

* - fix

* - fix

* - fix

* fix squueze_transpose (#47911)
Co-authored-by: NJacek Czaja <jacek.czaja@intel.com>

9e2ba9b9

PaddlePaddle / Paddle 大约 2 年 前同步成功

PaddlePaddle / Paddle
大约 2 年前同步成功