提交 · 73473ac26c03395cfaaa3087c4fa82ad318266e5 · PaddlePaddle / Paddle

09 4月, 2023 1 次提交
- register bf16 for c ops (#52641) · 73473ac2
  由 shaojie_wang 提交于 4月 08, 2023
  
  73473ac2
07 4月, 2023 1 次提交
- Z
  
  modify cmake file for cuda11.8 compile (#49020) (#52481) · 9431bae1
  由 zhaoyingli 提交于 4月 07, 2023
  
  9431bae1
03 4月, 2023 1 次提交
- Z
  
  make micro bsz configurable (#52447) · 722f880e
  由 zhaoyingli 提交于 4月 03, 2023
  
  722f880e
30 3月, 2023 1 次提交
- Y
  
  use int64 for c split (#52279) · 964497b5
  由 Yuang Liu 提交于 3月 30, 2023
  
  964497b5
28 3月, 2023 1 次提交
- L
  
  fix peak memory (#52175) · bd3b6adf
  由 LiYuRio 提交于 3月 28, 2023
  
  bd3b6adf
24 3月, 2023 1 次提交
- L
  
  optimize overlap between steps (#51974) · 81f4ef4f
  由 LiYuRio 提交于 3月 24, 2023
  
  81f4ef4f
20 3月, 2023 1 次提交
- L
  
  Cherry-pick fleet executor and auto parallel (#50071) · 92c2dcbd
  由 LiYuRio 提交于 3月 20, 2023
  
  92c2dcbd
09 3月, 2023 1 次提交
- J
  
  Extra Sync for Tensor Parallel (#50637) · 4bacf2ab
  由 JZ-LIANG 提交于 3月 09, 2023
  
  4bacf2ab
17 2月, 2023 1 次提交
- W
  
  Add rpc ops to fetch data from remote service (#50220) · 9025fddd
  由 Wen Sun 提交于 2月 17, 2023
  
  9025fddd
13 1月, 2023 2 次提交
- X
  
  fix_arg_release24 (#49771) · 0699afb1
  由 xiaoxiaohehe001 提交于 1月 13, 2023
  
  0699afb1
- Y
  fix fc kernel diff (#49781) · 01c26ab2
  由 Yuanle Liu 提交于 1月 13, 2023
```
* fix fc kernel diff

* disable fc_elementwise_layernorm_fuse_pass
```
  01c26ab2
12 1月, 2023 1 次提交
- X
  
  fix_split_infermeta (#49745) · 8a934047
  由 xiaoxiaohehe001 提交于 1月 12, 2023
  
  8a934047
09 1月, 2023 1 次提交
- H
  
  fix bugs of paddle.multiplex API (#49368) (#49642) · 6d2d8e50
  由 Haohongxiang 提交于 1月 09, 2023
  
  6d2d8e50
04 1月, 2023 2 次提交
- Y
  [Cherry-pick][Paddle Inference] fix mixed precision diff (#49477) · 1d25c663
  由 Yuanle Liu 提交于 1月 04, 2023
```
* disable scale op in amp pass

* Do not insert redundant cast op

* fix fused_fc_elementwise_layernorm kernel diff

* fix fc kerenl diff
```
  1d25c663
- Y
  [Cherry-pick] add condition of skipif (#49407) · 7696ae02
  由 YUNSHEN XIE 提交于 1月 04, 2023
```
* resolve conflict

* fix format error
```
  7696ae02
03 1月, 2023 2 次提交
- X
  [Cherry pick] fix fold for big bs (#49491) · 2a438b0a
  由 xiaoting 提交于 1月 03, 2023
```
* fix fold for large bs

* fix fold for large bs

* fix pre-commit
```
  2a438b0a
- F
  cherry-pick:Some version of TensorRT don't support qkv_plugin (#49425) · d7855fe8
  由 feng_shuai 提交于 1月 03, 2023
```
* cherry-pick:Some version of TensorRT don't support qkv_plugin

* cherry-pick:support coverage CI
```
  d7855fe8
30 12月, 2022 1 次提交

[MLU] cherry-pick from develop to release/2.4 (#48313) · 6e154fc6

由 Chenxiao Niu 提交于 12月 30, 2022

* [MLU] fix compute error of dropout op (#45923)

* [MLU] add mergedAdam kernel. (#45965)

* [MLU] add int64 support for mlu one_hot_v2 (#46313)

* [MLU] fix profiler compile failure (#46208)

* [MLU] add barrier_op kernel. (#46417)

* [MLU] fluid: add mluop (#46429)

* [MLU] add huber_loss kernel. (#46455)

* [MLU] add mlu kernel for add_reduce_max_grad (#45651)
Co-authored-by: Nliupeiyu <liupeiyu@cambricon.com>

* [MLU] add_fluid_mluop_yolo_box (#46573)

* [MLU] fix phi::Tensor compile error of mlu. (#46649)

* [MLU] add fluid MLUOps prior_box (#46585)

* [MLU] fix cmake error (#46772)

* [MLU]fix unittest of sync_bn (#46797)

* [MLU] add masterparam support for mlu adamw. (#46804)

* [MLU] add int64 support for allgather. (#46830)

* [MLU] fix compile error & add mlu blacklist function. (#47439)

* [MLU] fix softmax_with_cross_entropy failed in 370-X8.

* [MLU] fix cncl stuck caused by multiple initializations.

* [MLU] fix code style check.
Co-authored-by: Nqipengh <huangqipeng@cambricon.com>
Co-authored-by: Ncifar10 <41565156+cifar10@users.noreply.github.com>
Co-authored-by: Lux et Veritas <1004239791@qq.com>
Co-authored-by: Nliupeiyu <liupeiyu@cambricon.com>
Co-authored-by: Nronnywang <ronny1996@163.com>

6e154fc6

29 12月, 2022 2 次提交
- [cherry-pick]fix bug of UT test_version, test=document_fix (#49401) · 96e974a0
  由 zhouweiwei2014 提交于 12月 29, 2022
  
  96e974a0
- Y
  [Cherry-pick]Move sum op to PHI && Fix MetaTensor's bug when run infermeta (#49342) · 8015fbd6
  由 YuanRisheng 提交于 12月 29, 2022
```
* cherry-pick 45860

* [BUG FIX]Fix MetaTensor's bug when run infermeta (#46265)

* fix sum bug

* fix ci bugs

* fix ci bugs

* update code according comment
```
  8015fbd6
28 12月, 2022 1 次提交
- H
  [Cherry-pick] Fix CUDA11.8 Unittest Accuracy (#49374) · 8aa5be90
  由 Huihuang Zheng 提交于 12月 28, 2022
```
Fix CUDA11.8 Unittest Accuracy
```
  8aa5be90
27 12月, 2022 2 次提交

Y

update jetson ampere sm (#49364) · b5fdd175
由 Yuanle Liu 提交于 12月 27, 2022

b5fdd175

[Cherry-pick] Fix custom operator backward=None (#48656) (#48715) · 39eb77a6

由 HongyuJia 提交于 12月 27, 2022

* [Release2.4] Revert python link prs (#48573)

* Revert "Fix mac link python (#48017)"

This reverts commit 3fa7a736.

* Revert "[Cherry-pick] Fix python link error (#47811)"

This reverts commit ff642c68.

* Update config.go

* fix custom operator backward=None (#48656)

* [Custom Extension] Fix custom double_grad backward=None (#49224)

* fix custom double_grad backward=None

* fix custom_relu.cu bug && polish testcase of double_grad

* remove old dynamic graph test

* add import fluid

* add import fluid
Co-authored-by: NChen Weihang <chenweihang@baidu.com>

39eb77a6

22 12月, 2022 3 次提交

G

fix unittest in post training quantization (#49257) · 5d29a5bf
由 Guanghua Yu 提交于 12月 22, 2022

5d29a5bf

Fix mixed precision bug (#49239) · 11c7f570

由 Yuanle Liu 提交于 12月 22, 2022

* [Release2.4] Revert python link prs (#48573)

* Revert "Fix mac link python (#48017)"

This reverts commit 3fa7a736.

* Revert "[Cherry-pick] Fix python link error (#47811)"

This reverts commit ff642c68.

* Update config.go

* fix mixed precision inference
Co-authored-by: NChen Weihang <chenweihang@baidu.com>

11c7f570

L

[Docs]update readme; test=document_fix (#49246) · 612bdb17
由 Ligoml 提交于 12月 22, 2022

612bdb17

21 12月, 2022 2 次提交
- A
  
  fix unittests (#49203) (#49210) · 7c36b887
  由 Aganlengzi 提交于 12月 21, 2022
  
  7c36b887
- Z
  
  cherry-pick #75b734 (#49201) · fb19648a
  由 zhangkaihuo 提交于 12月 21, 2022
  
  fb19648a
20 12月, 2022 1 次提交
- S
  Fix nullptr to TestFuseGemmEpilogueReluBWDFP* (#48997) (#49090) · cdab3a44
  由 ShenLiang 提交于 12月 20, 2022
```
Co-authored-by: NMing-Xu Huang <mingh@nvidia.com>
```
  cdab3a44
19 12月, 2022 1 次提交

[cherry-pick][Inference] support mixed precision inference (#49077) · ddcd1b61

由 Yuanle Liu 提交于 12月 19, 2022

* [Release2.4] Revert python link prs (#48573)

* Revert "Fix mac link python (#48017)"

This reverts commit 3fa7a736.

* Revert "[Cherry-pick] Fix python link error (#47811)"

This reverts commit ff642c68.

* Update config.go

* [Paddle Inference] Add float_to_half_pass to support  inference with mixed precision (#47993)

* [Inference] optimize some code and fix some bug (#48780)

* clean ir_pass_manager and fix map_depthwise_conv_to_conv_pass

* fix unitest timeout

* [Paddle Inference] clean unused code  (#48392)

* fix

* update

* update
Co-authored-by: NChen Weihang <chenweihang@baidu.com>

ddcd1b61

29 11月, 2022 1 次提交

[cherry-pick] updating mul and matmul with set_mem_desc and fix... · 9e2ba9b9

由 yeliang2258 提交于 11月 29, 2022

[cherry-pick] updating mul and matmul with set_mem_desc and fix squeeze_transpose for MKLDNN (#47951)

* Fix slice bugs in MKLDNN when input dims are zeros (#46671)

* fix slice bugs

* fix

* update code

* fix

* update code

* updating mul and matmul with set_mem_desc (#45624)

* - mul & matmul changes

- fix

- bs16 correction of strides

* - cosmetic fixes

* - lint

* - fix

* - fix

* - format -> mem_desc

* - fix

* - fix

* - fix

* - fix

* - fix

* fix squueze_transpose (#47911)
Co-authored-by: NJacek Czaja <jacek.czaja@intel.com>

9e2ba9b9

28 11月, 2022 1 次提交

Cherrypick NV fixes to release/2.4 (#48263) · 7a0b8625

由 zlsh80826 提交于 11月 28, 2022

* Reduce squeeze2_matmul_fuse_pass, flattent tests time (#47098)

* Add missing fp32 config and reduce the testing combination

* Reduce trt matmul pass test max examples

* Loose TRT fp16 tests tolerance (#47100)

* Loose TRT half test tolerance to 1e-3 (#47101)

* Loose TRT half test tolerance to 1e-3 (#47106)

* Update distributed_strategy.proto (#46531)

* Close popen pipe after used (#47053)

* Add launch_bounds (#47285)

* Fix TRT UT failures (#47488)

* Format cherry-picked commits

* CudnnNormConvolution is no longer supported on NVIDIA Hopper GPUs (#48203)

* Skip tests that use fused_ops on H100

* Add error message to FusedOps on H100
Co-authored-by: NShijie <505749828@qq.com>
Co-authored-by: NLeo Chen <39020268+leo0519@users.noreply.github.com>
Co-authored-by: NTian Zheng <tizheng@nvidia.com>

7a0b8625

25 11月, 2022 2 次提交
- Z
  Fix wrong eigen header include in data_type.h (#48157) (#48260) · a2f61fef
  由 zyfncg 提交于 11月 25, 2022
```
* Fix wrong eigen header include

* fix compile bug
```
  a2f61fef
- W
  
  update (#48350) · b9b7f009
  由 Wilber 提交于 11月 25, 2022
  
  b9b7f009
24 11月, 2022 1 次提交
- U
  [cherry-pick2.4]en-docs warning&error fix (#48332) · 1490aaa9
  由 ustiniankw 提交于 11月 24, 2022
```
* fixdocs, test=document_fix

* fixdocs, test=document_fix
```
  1490aaa9
16 11月, 2022 1 次提交
- W
  Fix mac link python (#48017) · 3fa7a736
  由 wanghuancoder 提交于 11月 16, 2022
```
* finx mac link python

* refine
```
  3fa7a736
11 11月, 2022 2 次提交
- Y
  Fix slice bugs in MKLDNN when input dims are zeros (#46671) (#47887) · 5033b6c2
  由 yeliang2258 提交于 11月 11, 2022
```
* fix slice bugs

* fix

* update code

* fix

* update code
```
  5033b6c2
- H
  
  rename fw_bw func name of interleave pp (#47571) (#47862) · 4465ba27
  由 Haohongxiang 提交于 11月 11, 2022
  
  4465ba27
10 11月, 2022 2 次提交
- R
  Fuse multi transformer layer pass (#47541) (#47830) · 3a6cc57c
  由 RichardWooSJTU 提交于 11月 10, 2022
```
* add fuse_multi_transformer_layer_pass
```
  3a6cc57c
- W
  【cherry-pick】update Recompute doc (#47784) · 2e9e65d8
  由 wuhuachaocoding 提交于 11月 10, 2022
```
* cherry-pick recompute doc update.

* update.
```
  2e9e65d8

PaddlePaddle / Paddle 1 年多 前同步成功

PaddlePaddle / Paddle
1 年多前同步成功