提交 · 9783e887a728c671ee36654a25d3195639ee863c · Crayon鑫 / Paddle

21 6月, 2022 1 次提交
- G
  [cherry pick #43088 #40664] Add float16 to fake quantize/dequantize OP (#43689) · 9783e887
  由 Guanghua Yu 提交于 6月 21, 2022
```
* cherry pick #43088 #40664

* fix clang format
```
  9783e887
20 6月, 2022 5 次提交
- [cherry-pick]to Release/2.3,modify scale op xpu unittest (#43657) · 6262efb5
  由 z8hanghuan 提交于 6月 20, 2022
```
* modify xpu.cmake,*test=kunlun (#41832)

* modify xpu.cmake,*test=kunlun

* modify xpu.cmake,*test=kunlun

* modify xpu.cmake,*test=kunlun

* modify xpu.cmake,*test=kunlun

* support bilstm,*test=kunlun

* [cherry-pick]support multi_layer of bilstm,*test=kunlun

* [cherry-pick]refactor sum unit test,*test=kunlun (#43561)
```
  6262efb5
- X
  [Cherry pick] Einsum memory optimization PR #43397 (#43554) · 638b69dc
  由 xiongkun 提交于 6月 20, 2022
```
* cherry pick from #43397

* fix code
```
  638b69dc
- S
  
  fix unittest (#43609) (#43617) · 68d5c12b
  由 Shang Zhizhou 提交于 6月 20, 2022
  
  68d5c12b
- Z
  
  place all save/load path into temporary directory (#43652) · a5ccc713
  由 zhaoyingli 提交于 6月 20, 2022
  
  a5ccc713
- Z
  [Cherry-Pick] place all save/load path into temporary directory (#43316) (#43651) · 0f16ccf5
  由 zhaoyingli 提交于 6月 20, 2022
```
* place all save/load path into temporary directory

* rm no need unittest
```
  0f16ccf5
17 6月, 2022 2 次提交

Y

cherry pick 43581 (#43596) · 2eb60ddb
由 YuanRisheng 提交于 6月 17, 2022

2eb60ddb

[cherry-pick 2.3] Cherry parallel fused transformer api (#43505) · 19b87aec

由 WangXi 提交于 6月 17, 2022

* Rename dropout is test (#43098)

* replace dropout_is_test with is_test.
* improve atol on a100.

* fused_attention fused_feedforward api support Model Tensor Parallel (#42985)

* fix is_test bug in fused_feedforward. (#43508)
Co-authored-by: NLi Min <11663212+limin2021@users.noreply.github.com>

19b87aec

16 6月, 2022 4 次提交

[cherry pick] Unit test with tempfile to place the temporary files (#43522) · 1a660c8a

由 zhangbopd 提交于 6月 16, 2022

Use tempfile for unit test & custom op test to replace temporary files to ensure that all temporary files will be deleted normally after a single measurement, avoiding the usage of disk files.
The PR only involves single-test and op test modifications and does not affect existing functionality.
Release/2.3 branch modified in PR43521;

1a660c8a

Q
[Cherry-pick] Fix ut tempfile v23 (#43387) · 24843fcb
由 Qi Li 提交于 6月 16, 2022
```
* fix unit test temp file, test=develop (#43155)

* add cleanup code, test=develop (#43305)
```
24843fcb

[Cherry-pick] Fix numpy 1.20+ deprecation warnings (#43513) · 689e0999

由 Qi Li 提交于 6月 16, 2022

* Fix numpy 1.20+ deprecation warnings (#42929)

* Replace np.bool/np.bool8 with np.bool_

* Replace np.object with np.object_

* Replace np.complex with np.complex128

* Replace np.float with np.float64

* Replace np.int with np.int_

* Rerun pre-commit for newer pre-commit configuration

* Use builtin bool instead of np.bool_ based on the context

* fix mode dtype
Co-authored-by: Nzlsh80826 <rewang@nvidia.com>

689e0999

Z

cherry-pick adamw unittest (#43498) · 0cdde0b4
由 zhaoyingli 提交于 6月 16, 2022

0cdde0b4

14 6月, 2022 2 次提交

[ CherryPick ] Cherry pick for einsum optimization. (#43468) · 22e75d92

由 xiongkun 提交于 6月 14, 2022

* [EinsumOp] Polish forward logic and backward logic for optimize (#42603)

* change logic for optimize

* modifty

* merge

* change einsum_v2 as default and add new flags: FLAG_einsum_opt=1|0 (#43010)

* [EinsumOp] Make EinsumOp support bfloat16. (#43085)

* change einsum_v2 as default and add new flags: FLAG_einsum_opt=1|0

* make EInsumOP support bf16

* add unittest for BF16

* add condition for test_BF16

* fix bugs

* fix

* change the backward api to fit einsum op

22e75d92

Use tempfile to place all the temporary files. (#43392) · afd0c1db

由 freeliuzc 提交于 6月 14, 2022

使用 tempfile 替换临时文件，保证在单测结束后，所有临时文件都会被正常的删除，避免占用磁盘文件。
此 PR 仅涉及单测修改，不影响现有功能。
develop 分支修改在 PR 43376

afd0c1db

09 6月, 2022 1 次提交
- G
  
  Modify quantization use tempfile to place the temporary files (#43281) · f4e09397
  由 Guanghua Yu 提交于 6月 09, 2022
  
  f4e09397
30 5月, 2022 1 次提交
- W
  [Dy2St]Fix cond_block_grad error when handle no need grad vras (#43034) (#43084) · e6e85b35
  由 WangZhen 提交于 5月 30, 2022
```
* Fix cond_block_grad error when handle no need grad vras

* Add comment and UT
```
  e6e85b35
26 5月, 2022 1 次提交
- S
  make some test run with old executor in specified windows server (#42777) (#42981) · 7a223585
  由 Sing_chan 提交于 5月 26, 2022
```
cherry-pick PR #42777
```
  7a223585
19 5月, 2022 1 次提交
- A
  [Dy2Stat]Modify all jit.save path into tempfile under dygraph_to_static directory (#42842) (#42860) · 84840481
  由 Aurelius84 提交于 5月 19, 2022
```
* [Dy2Stat]Modify all jit.save path into tempfile

* [Dy2Stat]Modify all jit.save path into tempfile
```
  84840481
10 5月, 2022 1 次提交

[cherry-pick][MLU] support add callback to stream and profiler (#42115) · 25124d7f

由 fwenguang 提交于 5月 10, 2022

* [MLU] add mlu new profiler (#41138)

* [MLU] add mlu new profiler

* fix format

* [MLU] support add callback to stream (#41831)

* [MLU] add gather mlu kernel (#41969)

* [MLU] add mlu activation kernels (#41751)

25124d7f

09 5月, 2022 1 次提交

[Cherry-pick][IPU] merge recent changes (#42078) (#42582) · 1f9b60df

由 Allen Guo 提交于 5月 09, 2022

    add class NameScopeHelper for adding namescope info
    添加更多 种类优化器状态的映射
    为 IpuStrategy 添加 compilation_progress_logger option 用于输出 编译进度
    部分代码清理和杂项优化

1f9b60df

07 5月, 2022 2 次提交
- W
  
  remove the test case for the matmul_v2_mkldnn (#42530) · 54ef3d56
  由 wawltor 提交于 5月 07, 2022
  
  54ef3d56
- R
  [cherry-pick] Fix UT timeout problem for cuda_managed_memory_test and test_tensordot (#42492) · c9d156b1
  由 Ruibiao Chen 提交于 5月 07, 2022
```
* Reduce time variation for cuda_managed_memory_test (#42458)

* Disable standalone executor for test_tensordot (#42476)
```
  c9d156b1
06 5月, 2022 1 次提交
- L
  [cherry-pick] fix wrong place in ut (#42488) · 35ed11f3
  由 Leo Chen 提交于 5月 06, 2022
```
* fix wrong place

* skip bf16 test if not supported (#42503)
```
  35ed11f3
05 5月, 2022 2 次提交
- W
  
  fix unittest of conv2d due to V100 do not support bfloat16 (#42496) · 71d3b06c
  由 wangxinxin08 提交于 5月 05, 2022
  
  71d3b06c
- W
  
  fix the v100 cuda11.2 matmul_v2 and elementwise_div bug (#42479) · e052fde7
  由 wawltor 提交于 5月 05, 2022
  
  e052fde7
03 5月, 2022 1 次提交
- H
  Hotfix Release 2.3 Bug for CUDA 11.2 (#42438) · 713d5a4b
  由 Huihuang Zheng 提交于 5月 03, 2022
```
* Fix Release 2.3 Bug

* Fix format
```
  713d5a4b
30 4月, 2022 4 次提交
- A
  [Dy2Stat]Fix losting pre/post hook from outermost layer while jit.save (#42273) (#42388) · 16ef2b2e
  由 Aurelius84 提交于 4月 30, 2022
```
* [Dy2Stat]Fix losting pre/post hook from outermost layer while jit.save

* fix kwargs

* fix unittest
```
  16ef2b2e
- W
  
  [Eager] Support test_diff_op switch to eager mode (#42360) (#42392) · 1e3d2e4a
  由 Weilong Wu 提交于 4月 30, 2022
  
  1e3d2e4a
- X
  Make einsum_v2 support multi-operands (#42327) (#42397) · 34352fcd
  由 xiongkun 提交于 4月 30, 2022
```
* Extend python einsum interface to make einsum_v2 support multi-operands and switch it to default.

* add opt_einsum dependence

* add yaml and support eager model

* fix by code review
```
  34352fcd
- R2.3/fix pad3d infer shape (#42414) · 2dce1e88
  由 littletomatodonkey 提交于 4月 30, 2022
```
* fix pad3d infer shape

* fix pad3d

* fix pad default value

* fix order

* add unit test

* fix unittest for ci coverage

* add ndhwc check
```
  2dce1e88
29 4月, 2022 2 次提交

[cherry-pick 2.3] Add fused_multi_transformer op to optimize transformer... · 50bfe420

由 WangXi 提交于 4月 29, 2022

[cherry-pick 2.3] Add fused_multi_transformer op to optimize transformer generation performance (#42311)

* Add fused_multi_transformer op to optimize transformer generation performance (#41814)

* fix fused_multi_transformer compile failed in cuda arch < sm53 (#42315)

* fix ci timeout

50bfe420

Z
[cherry-pick] Fix bug of building InferMetaContext (#42211) (#42399) · 765fbb59
由 zyfncg 提交于 4月 29, 2022
```
* fix bug of building InferMetaContext (#42211)

* add unitest
```
765fbb59

28 4月, 2022 2 次提交

Add C++ EinsumOp which support 2 operands einsum. (#42105) (#42357) · d04a68d3

由 xiongkun 提交于 4月 28, 2022

* full api fix

* when out is None, go old dygraph mode

* by static check

* first version: support 2-inputs forwards. TODO: 1. backward  2. BroadCast  3. MultiVariable

* time out -> 120

d04a68d3

[cherry-pick] implement autotune python API(42299) (#42301) · b37e626e

由 Zhang Ting 提交于 4月 28, 2022

* implement autotune python API

* fix doc

* fix windows error

* fix doc and enable auto-tuning when config is None

* fix windows error

* fix doc

b37e626e

27 4月, 2022 3 次提交
- P
  
  fix format · 880c2a94
  由 pangyoki 提交于 4月 26, 2022
  
  880c2a94
- P
  
  solve conflict · eade1fd9
  由 pangyoki 提交于 4月 25, 2022
  
  eade1fd9
- W
  [Eager] Remove retain_grad_flag in accumulation_nade, add is_new_grad args in... · 56b93800
  由 Weilong Wu 提交于 4月 27, 2022
```
[Eager] Remove retain_grad_flag in accumulation_nade, add is_new_grad args in operator (#42240) (#42290)
```
  56b93800
26 4月, 2022 3 次提交
- W
  
  [Eager] Support numpy.ndarry in CastNumpy2Scalar (#42136) (#42213) · 983fcb56
  由 Weilong Wu 提交于 4月 26, 2022
  
  983fcb56
- fix python3.10 compile bug on windows (#42140) (#42180) · 42297995
  由 zhouweiwei2014 提交于 4月 26, 2022
```
cherry-pick #42140
```
  42297995
- W
  [Eager] Support div(scalar) in eager mode (#42148) (#42214) · a887ffd0
  由 Weilong Wu 提交于 4月 26, 2022
```
* [Eager] Support div scalar in eager mode

* Updated and remove debug logs

* Remove list, use 'or' directly

* Remove useless statement
```
  a887ffd0

Crayon鑫 / Paddle 与 Fork 源项目一致

Crayon鑫 / Paddle
与 Fork 源项目一致