提交 · 22e75d92cb04ede350e84be9a774b1b528f650b1 · Crayon鑫 / Paddle

14 6月, 2022 2 次提交

[ CherryPick ] Cherry pick for einsum optimization. (#43468) · 22e75d92

由 xiongkun 提交于 6月 14, 2022

* [EinsumOp] Polish forward logic and backward logic for optimize (#42603)

* change logic for optimize

* modifty

* merge

* change einsum_v2 as default and add new flags: FLAG_einsum_opt=1|0 (#43010)

* [EinsumOp] Make EinsumOp support bfloat16. (#43085)

* change einsum_v2 as default and add new flags: FLAG_einsum_opt=1|0

* make EInsumOP support bf16

* add unittest for BF16

* add condition for test_BF16

* fix bugs

* fix

* change the backward api to fit einsum op

22e75d92

Use tempfile to place all the temporary files. (#43392) · afd0c1db

由 freeliuzc 提交于 6月 14, 2022

使用 tempfile 替换临时文件，保证在单测结束后，所有临时文件都会被正常的删除，避免占用磁盘文件。
此 PR 仅涉及单测修改，不影响现有功能。
develop 分支修改在 PR 43376

afd0c1db

13 6月, 2022 1 次提交
- T
  2.3 del test message (#43404) · ed859054
  由 tianshuo78520a 提交于 6月 13, 2022
```
删除无用信息
```
  ed859054
09 6月, 2022 3 次提交
- G
  cherry pick #42255 (fuse conv + bn in QAT) and #42378 (support skip_op_list in PTQ) (#43301) · 0a00fc4e
  由 Guanghua Yu 提交于 6月 09, 2022
```
* support fuse conv and bn in QAT (#42255)

* support skip_op_list in PostTrainingQuantization (#42378)

* fix unittest
```
  0a00fc4e
- G
  
  Modify quantization use tempfile to place the temporary files (#43281) · f4e09397
  由 Guanghua Yu 提交于 6月 09, 2022
  
  f4e09397
- Z
  
  disable lite gpu (#43178) · 36980306
  由 zhupengyang 提交于 6月 09, 2022
  
  36980306
08 6月, 2022 4 次提交
- N
  Replace ReduceAmax/Amax.part.cu with KP (#43202) (#43263) · e161979e
  由 niuliling123 提交于 6月 08, 2022
```
Reduce amax/amin frobenius_norm_kerne原始实现为Eigen实现，文件编译时间较长，因此本PR将其替换为KP实现
删除DefaultElementwiseOperator中重复功能支持，减少elementwise_double_grad OP编译时间
```
  e161979e
- T
  Del whl check for release/2.3 (#43288) · 8f127681
  由 tianshuo78520a 提交于 6月 08, 2022
```
删除在2.3 对比whl包大小。
```
  8f127681
- J
  
  updated paddle_bfloat to v0.1.7 · df1d4645
  由 jakpiase 提交于 5月 19, 2022
  
  df1d4645
- H
  Resolve protobuf of ORT Backend conflict (#43275) · c2804390
  由 heliqi 提交于 6月 07, 2022
```
解决onnxruntime后端依赖的protobuf跟框架或外部protobuf版本冲突问题
```
  c2804390
07 6月, 2022 3 次提交
- Z
  
  fix the problem of slice infer shape (#42568) (#43246) · f1b4e4d5
  由 zyfncg 提交于 6月 07, 2022
  
  f1b4e4d5
- X
  
  fix memory leakage (#43141) (#43220) · e09803c5
  由 xiongkun 提交于 6月 07, 2022
  
  e09803c5
- N
  [cherry-pick]Delete ElementwiseKernel in BroadcastKernel (#42779) (#43210) · 52ef8656
  由 niuliling123 提交于 6月 07, 2022
```
Delete ElementwiseKernel in BroadcastKernel
减少所有Broadcast中重复功能调用，同时减少编译时间和问题体积
```
  52ef8656
06 6月, 2022 1 次提交

cherry-pick 42645 (#43205) · 835a1888

由 niuliling123 提交于 6月 06, 2022

删除Broadcast function中rank例化以及Elementwise调用，降低编译时间。
从develop分支中的#42645 PR修改而来，由于develop分支与release分支相差较大，无法实现cherry-pick，因此针对release2.3重新提交PR.
Broadcast中关于rank的例化会导致底层模板展开较多，造成reduce_sum_grad_kernel.cu.o文件体积过大，修改后可以降低.o体积及编译时间

835a1888

31 5月, 2022 1 次提交

Del check size (#43113) · 40a7e0ad

由 tianshuo78520a 提交于 5月 31, 2022

删除判断build目录大小和预测库大小检查功能。该功能是和develop比较，会存在差异，在release任务中取消判断

40a7e0ad

30 5月, 2022 2 次提交
- W
  [Dy2St]Fix cond_block_grad error when handle no need grad vras (#43034) (#43084) · e6e85b35
  由 WangZhen 提交于 5月 30, 2022
```
* Fix cond_block_grad error when handle no need grad vras

* Add comment and UT
```
  e6e85b35
- W
  [Paddle-Inference] fix_multiheadpass_int8 (#43020) · 72880279
  由 Wangzheee 提交于 5月 30, 2022
```
* fix_multi_int8 (#42977)

* cherry-pick fix_multihead_int8
```
  72880279
27 5月, 2022 4 次提交
- T
  
  test=document_fix · aedd4592
  由 tianshuo78520a 提交于 5月 27, 2022
  
  aedd4592
- T
  
  test=document_fix · 67da108a
  由 tianshuo78520a 提交于 5月 27, 2022
  
  67da108a
- T
  
  cherry-pick · 8bff7f9b
  由 tianshuo78520a 提交于 5月 27, 2022
  
  8bff7f9b
- T
  
  Test release/2.3 ci · e7871ff3
  由 tianshuo78520a 提交于 4月 06, 2022
  
  e7871ff3
26 5月, 2022 2 次提交
- S
  make some test run with old executor in specified windows server (#42777) (#42981) · 7a223585
  由 Sing_chan 提交于 5月 26, 2022
```
cherry-pick PR #42777
```
  7a223585
- C
  
  polish kernel type str (#42791) (#42931) · b5766fbf
  由 Chen Weihang 提交于 5月 26, 2022
  
  b5766fbf
23 5月, 2022 2 次提交
- O
  Update metrics.py · d5b6eec2
  由 onecatcn 提交于 5月 19, 2022
```
the doc was editted based on the discussion in the issue:
INT32 Failed on paddle.metric.accuracy: https://github.com/PaddlePaddle/Paddle/issues/42845
```
  d5b6eec2
- S
  【CI】run all demo ci before exit in windows (#42700) (#42897) · 2300d45f
  由 Sing_chan 提交于 5月 23, 2022
```
cherry-pick PR #42700
```
  2300d45f
19 5月, 2022 1 次提交
- A
  [Dy2Stat]Modify all jit.save path into tempfile under dygraph_to_static directory (#42842) (#42860) · 84840481
  由 Aurelius84 提交于 5月 19, 2022
```
* [Dy2Stat]Modify all jit.save path into tempfile

* [Dy2Stat]Modify all jit.save path into tempfile
```
  84840481
17 5月, 2022 2 次提交
- C
  
  fix trace op record event error (#42775) (#42789) · af79273d
  由 Chen Weihang 提交于 5月 17, 2022
  
  af79273d
- C
  put_record_event_in_python_on_timeline_python (#42555) (#42790) · a40e60f7
  由 chenjian 提交于 5月 17, 2022
```
* put_record_event_in_python_on_timeline_python

* fix
```
  a40e60f7
16 5月, 2022 1 次提交
- W
  fix sample code error of paddle.lerp, test=document_fix (#42753) · 07029e0c
  由 wuhuanzhou 提交于 5月 16, 2022
```
修复paddle.lerp中示例代码错误。
```
  07029e0c
11 5月, 2022 1 次提交
- A
  
  [Eager]Fix EagerTensor _copy_to memory overlap problem (#42668) (#42686) · d0e733dd
  由 Aurelius84 提交于 5月 11, 2022
  
  d0e733dd
10 5月, 2022 4 次提交
- J
  pdnode_compare (#42597) (#42633) · 403b503f
  由 JingZhuangzhuang 提交于 5月 10, 2022
```
* pdnode_compare

* panode compare

* pdnode_compare
```
  403b503f
- F
  [cherry-pick][MLU] support add callback to stream and profiler (#42115) · 25124d7f
  由 fwenguang 提交于 5月 10, 2022
```
* [MLU] add mlu new profiler (#41138)

* [MLU] add mlu new profiler

* fix format

* [MLU] support add callback to stream (#41831)

* [MLU] add gather mlu kernel (#41969)

* [MLU] add mlu activation kernels (#41751)
```
  25124d7f
- A
  set custom_nll_loss_op attr ignoreIndex to str (#42596) · 6c935e1d
  由 Allen Guo 提交于 5月 10, 2022
```
set attr ignoreIndex type to string for custom_nllloss_op

部分 cheery-pick of #42534
```
  6c935e1d
- Z
  
  fix bug of optional_tensor in amp logic (#42561) (#42577) · 37715dab
  由 zhangbo9674 提交于 5月 10, 2022
  
  37715dab
09 5月, 2022 1 次提交

[Cherry-pick][IPU] merge recent changes (#42078) (#42582) · 1f9b60df

由 Allen Guo 提交于 5月 09, 2022

    add class NameScopeHelper for adding namescope info
    添加更多 种类优化器状态的映射
    为 IpuStrategy 添加 compilation_progress_logger option 用于输出 编译进度
    部分代码清理和杂项优化

1f9b60df

07 5月, 2022 3 次提交
- W
  
  remove the test case for the matmul_v2_mkldnn (#42530) · 54ef3d56
  由 wawltor 提交于 5月 07, 2022
  
  54ef3d56
- F
  Reduce the number of threads per block of deformable_psroi_pooling to solve... · 44271ece
  由 FlyingQianMM 提交于 5月 07, 2022
```
Reduce the number of threads per block of deformable_psroi_pooling to solve the bug where too many resources requested for launch (PaddlePaddle#42531) (#42533)
```
  44271ece
- R
  [cherry-pick] Fix UT timeout problem for cuda_managed_memory_test and test_tensordot (#42492) · c9d156b1
  由 Ruibiao Chen 提交于 5月 07, 2022
```
* Reduce time variation for cuda_managed_memory_test (#42458)

* Disable standalone executor for test_tensordot (#42476)
```
  c9d156b1
06 5月, 2022 2 次提交
- L
  [cherry-pick] fix wrong place in ut (#42488) · 35ed11f3
  由 Leo Chen 提交于 5月 06, 2022
```
* fix wrong place

* skip bf16 test if not supported (#42503)
```
  35ed11f3
- W
  Fix the race condition in cumsum operator (#42205) (#42500) · 58f40144
  由 wawltor 提交于 5月 06, 2022
```
* Fix the race condition in cumsum operator

* Optimize cumsum operator
Co-authored-by: NLeo Chen <39020268+leo0519@users.noreply.github.com>
```
  58f40144

Crayon鑫 / Paddle 与 Fork 源项目一致

Crayon鑫 / Paddle
与 Fork 源项目一致