提交 · 0c31579c1c0242e184fe2dc7f8e14f4949da62a7 · PaddlePaddle / Paddle

13 10月, 2021 3 次提交

由 limingshu 提交于 10月 13, 2021

* A leap of try for cudaLaunchCooperativeKernel

* fix bugs

* Totally replace the lar cuda kernel

* Fix bugs

* a test for lars merge

* Adding las_op_momentum infer_shape

* Fix codes

* use avg_numel instead of max_numel to acquire grid num

* modify unittest files about lars op

* Finally converge when merged-lars works

* fix ctest files

* add merged_operation kernel when cuda version is older than 11

* Fix code style

* fix ctest failure

* fix error

* fix all ctest error and change lars compute code of cpu

* fix bugs on v100.

* revert python modififation about lars

* revert python modification codes

0c31579c

J
Implemented LRU based cache clearing (#36290) · bf748f24
由 Jacek Czaja 提交于 10月 13, 2021
```
- Lint

- Merge with develop

- lint
```
bf748f24

[New Feature] Support triple grad in Paddle (#36187) · 2c44ee7e

由 Jiabin Yang 提交于 10月 13, 2021

* native commit for triple grad of sigmod

* Updated unittests files

* init functional jacobian api

* Updated trible_test func

* Updated gradient_checker & test_script

* finish test with dtype float32

* add float64 test case

* polish code

* use atol=1e-5 with dtype float64

* fix for ci

* set timeout for test_jacobian

* fix dygraph grad to support high differential

* polish API docstring

* Updated gradient checker and some related files

* fix double grad strip error for high differential

* fix double grad strip error for high differential

* Add Sigmoid triple grad tests

* fix dygraph double grad dtype error when calling for high differential senario

* Updated triple grad teses func

* Use np.random to initialize ddx

* Updated triple_grad_check func

* add todo for gradient checker and refine some comments

* remove additional code

* add test for warnging in backward.py

* format python code
Co-authored-by: Nveyron95 <veyron_wu@163.com>
Co-authored-by: Nlevi131 <limaolin01@baidu.com>

2c44ee7e

12 10月, 2021 5 次提交
- Z
  
  Change the input param of fusion op interface from pointer to tensor (#36349) · 3e2dec5b
  由 Zhang Zheng 提交于 10月 12, 2021
  
  3e2dec5b
- A
  [NPU] concat supports dtype int64 for model deepfm (#36327) · 5f1eb839
  由 Aganlengzi 提交于 10月 12, 2021
```
* [NPU] modify for model deepfm

* [NPU] unit test delete precision control

* [NPU] add more unit test

* revert elementwise_mul related modification

* [NPU] add more unit tests for concat
```
  5f1eb839
- Q
  [NPU] fix elementwise_mul to support broadcast, test=develop (#36258) · 09778f46
  由 Qi Li 提交于 10月 12, 2021
```
* [NPU] fix elementwise_mul to support broadcast, test=develop

* remove debug files, test=develop

* add axis support, test=develop
```
  09778f46
- Q
  [NPU] add int64 kernel for slice, test=develop (#36328) · 8cc7146d
  由 Qi Li 提交于 10月 12, 2021
```
* [NPU] add int64 kernel for scale and slice, test=develop

* remove int64 for scale, test=develop
```
  8cc7146d
- A
  Fix stop_gradient in RunProgramOp (#36339) · 2a75b447
  由 Aurelius84 提交于 10月 12, 2021
```
* Fix stop_gradient in RunProgramOp

* fix reference
```
  2a75b447
11 10月, 2021 8 次提交
- J
  
  fix for matmul_v2 6D x 2D (#36342) · 339cb191
  由 jakpiase 提交于 10月 11, 2021
  
  339cb191
- L
  Add nn.functional.sparse_attention and some test cases, test=develop (#35757) · 85b77232
  由 Liu-xiandong 提交于 10月 11, 2021
```
Add paddle.nn.functional.sparse_attention API

    本个PR主要将sparse_attention功能在python层进行了一层封装，OP的主体代码见：#PR35676

    此外，对于封装的python 接口，增加了相应的单测。
```
  85b77232
- Z
  
  Add more tests and fix bugs for cudnn_norm_conv_test and cudnn_bn_and_relu_test (#36314) · a679fcbb
  由 Zhang Zheng 提交于 10月 11, 2021
  
  a679fcbb
- N
  Add functor_primitives.h for kernel primtive api (#36203) · 830debc2
  由 niuliling123 提交于 10月 11, 2021
```
* Add functor_primitives.h for kernel primtive api

* update

* move namespace kps

* subFunctor init_data

* delete InvalidArgumentError
```
  830debc2
- Q
  [NPU] fix matmul_v2 and utils.run_check, test=develop (#36164) · 7850f7ce
  由 Qi Li 提交于 10月 11, 2021
```
* [NPU] fix matmul_v2 and utils.run_check, test=develop

* remove debug files, test=develop

* fix install_check, test=develop

* fix doc, test=develop

* fix review comments, test=develop
```
  7850f7ce
- Q
  [NPU] fix set_value, test=develop (#36272) · 83541fd4
  由 Qi Li 提交于 10月 11, 2021
```
* [NPU] fix set_value, test=develop

* fix typo, test=develop

* fix typo, test=develop
```
  83541fd4
- Q
  
  [NPU] fix softmax_with_cross_entropy in dygraph, test=develop (#36297) · 11061325
  由 Qi Li 提交于 10月 11, 2021
  
  11061325
- X
  
  use unified external error message for cufft api (#36114) · 642aaa2e
  由 Xiaoxu Chen 提交于 10月 11, 2021
  
  642aaa2e
09 10月, 2021 3 次提交
- Z
  
  Implement Fused BN + Add + Relu with cudnnFusedOps API. (#35955) · 7e6c0cee
  由 Zhang Zheng 提交于 10月 09, 2021
  
  7e6c0cee
- Y
  
  Enhance OpTest for bfloat16. (#36079) · 91119271
  由 Yiqun Liu 提交于 10月 09, 2021
  
  91119271
- Z
  
  fill_diagonal op fix border cross caused by offset (#36212) · 62e41150
  由 zhiboniu 提交于 10月 09, 2021
  
  62e41150
08 10月, 2021 5 次提交
- J
  Fix for oneDNN conv op (#36284) · 57e8cbec
  由 jakpiase 提交于 10月 08, 2021
```
* fix for conv op

* Minor change
```
  57e8cbec
- Z
  Support CUDA Graph on ParallelExecutor (#36250) · f9591bb1
  由 Zeng Jinle 提交于 10月 08, 2021
```
* support CUDA Graph on PE

* add ut, fix CI compile

* reduce memory consumption

* fix CUDA 10 CI

* improve coverage

* improve python coverage
```
  f9591bb1
- Q
  [NPU] BatchNorm support layout of NCL and NLC, test=develop (#35668) · 7cb19f57
  由 Qi Li 提交于 10月 08, 2021
```
* [NPU] support NCL and NCL for BatchNorm, test=develop

* [NPU] remove debug files, test=develop

* update, test=develop
```
  7cb19f57
- A
  Added oneDNN BF16 relu (#36265) · 1bd9cfef
  由 arlesniak 提交于 10月 08, 2021
```
* Added oneDNN BF16 relu

* fixed typo

* refactored test, review fixes
```
  1bd9cfef
- Z
  
  fix cast cuda implementation (#36266) · 9814f895
  由 Zeng Jinle 提交于 10月 08, 2021
  
  9814f895
07 10月, 2021 1 次提交

[OneDNN] Conv op refactor. (#36252) · e9288340

由 Adam Osewski 提交于 10月 07, 2021

* Remove unused header.

* Use ConvMKLDNNHandlerT for conv2d INT8.

* Use absolute module path to import.

e9288340

05 10月, 2021 1 次提交

Added concat BF16/FP32 BWD OneDNN kernel (#35889) · dc4d5719

由 jakpiase 提交于 10月 05, 2021

* tmp

* added concat BF16/FP32 BWD oneDNN kernel

* minor change

* minor change

* fix for CI

* added formatting

* Reverted deleting static keyword

* added reviewers suggestions

* reverted deleting concat bf16 test file

* fixed concat tests

dc4d5719

30 9月, 2021 1 次提交

[NPU] modify transpose2 and index_select_grad kernels for model xlnet (#36214) · a66b9fba

由 Aganlengzi 提交于 9月 30, 2021

* [NPU] modify transpose2 and index_select_grad kernels for model xlnet

* add transpose2 int64_t unit test

* add more transpose2 unit tests

* update test_transpose_op_npu.py

a66b9fba

29 9月, 2021 7 次提交
- Z
  [npu] add box coder (#36171) · 83578cfa
  由 zhulei 提交于 9月 29, 2021
```
* [npu] add box coder

* [npu] add box coder
```
  83578cfa
- P
  
  fix bug of top_k npu op (#36175) · 2b8fd704
  由 pangyoki 提交于 9月 29, 2021
  
  2b8fd704
- Z
  [NPU] Add group norm (#35937) · c79de728
  由 zhulei 提交于 9月 29, 2021
```
* [NPU] Add group norm

* [NPU] Add group norm

* [NPU] Add group norm

* [NPU] Add group norm

* [NPU] Add group_norm op
```
  c79de728
- A
  [NPU] mod for model bert (#36165) · 7bddf2e8
  由 Aganlengzi 提交于 9月 29, 2021
```
* merge conflict of paddle_gtest_main.cc

* modify FLAGS_npu_precision_mode and default not to call aclSetCompileopt
```
  7bddf2e8
- Y
  
  Implement the grad and enhance the cache of norm_convolution fusion ops. (#36168) · 767050d9
  由 Yiqun Liu 提交于 9月 29, 2021
  
  767050d9
- L
  
  Add fused_dropout wrapper to ease use. (#36185) · 092d45c3
  由 Li Min 提交于 9月 29, 2021
  
  092d45c3
- R
  
  [ROCM] bugfix for bilinear_interp_v2_grad (#36160) · 5e1d0b5c
  由 ronnywang 提交于 9月 29, 2021
  
  5e1d0b5c
28 9月, 2021 5 次提交

L
Add sparse_attention api, test=develop (#35676) · 6b587e93
由 Liu-xiandong 提交于 9月 28, 2021
```
Add sparse_attention OPs, python api will be added in next pr
```
6b587e93

add API paddle.linalg.eig (#35674) · bc7e2b92

由 Lijunhui 提交于 9月 28, 2021

* Add paddle.linalg.eig op

* remove comments

* remove comments

* extend batch_size to the origin

* add real times complex functor & destroy the backward complex output bug

* terminate output diff when input real tensors

* correct tiny doc errors

* move functions from eig_helper to svd_helper and remove eig_helper

* remove tensor.Resize

* remove no longer used code

* use existing lapack functions

* reply review comments 21/27

* remove .cu as this op is only executed on CPU

* remove const_cast & add const in argument list for read-only references

* fix sample code error in CI

* remove template typename Tbase and more

* remove eig exposure in paddle.*

* add 'name=None' in eig python implementation

* handle the unittest

* try to solve the unittest

* solve CI coverage

* remove no longer used code

* polish API doc and more

* reply review comments

* polish unittest, commit plan B

* polish unittest

bc7e2b92

R

[ROCM] bugfix for arg_min_max (#36098) · 36791fdd
由 ronnywang 提交于 9月 28, 2021

36791fdd

[hybrid] seed and dropout op support force-cpu (#35820) · 58c8f6b3

由 xiayanming 提交于 9月 28, 2021

* [HIP] fix op not support AMD GPU bug, the flag PADDLE_WITH_ROCM is invalid

* [HIP] fix op not support AMD GPU bug, the flag PADDLE_WITH_ROCM is invalid

* [HIP] fix op not support AMD GPU bug

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] fix seed ci failed issue

* add AsExtra for force_cpu of seed op

58c8f6b3

G

fix bug of reduce_sum when src_dtype != dst_dtype and reduce_num == 1 (#36123) · d5268a6e
由 Guoxia Wang 提交于 9月 28, 2021

d5268a6e

27 9月, 2021 1 次提交

fix zero tensor for unique, unstack (#36021) · efd35384

由 Jiawei Wang 提交于 9月 27, 2021

* fix extra op for expand, expand_as, tile, unstack

* fix unique unstack dim 0

* Update expand_v2_op.cc

* fix unique_op format

efd35384

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功