提交 · 91119271584dbf6cefe86a170e078d245bf912e5 · Crayon鑫 / Paddle

09 10月, 2021 4 次提交
- Y
  
  Enhance OpTest for bfloat16. (#36079) · 91119271
  由 Yiqun Liu 提交于 10月 09, 2021
  
  91119271
- Z
  Add const for OpDesc::id() and VarDesc::id() (#36298) · cb620ca6
  由 Zeng Jinle 提交于 10月 09, 2021
```
* add const OpDesc id()

* add const for VarDesc::id()
```
  cb620ca6
- Z
  
  fill_diagonal op fix border cross caused by offset (#36212) · 62e41150
  由 zhiboniu 提交于 10月 09, 2021
  
  62e41150
- W
  C++ support register pass via PassDesc (#36095) · 2fd8deea
  由 wuhuanzhou 提交于 10月 09, 2021
```
支持C++开发注册GeneratePass，简化针对fusion等子图优化场景开发方式。
```
  2fd8deea
08 10月, 2021 6 次提交
- J
  Fix for oneDNN conv op (#36284) · 57e8cbec
  由 jakpiase 提交于 10月 08, 2021
```
* fix for conv op

* Minor change
```
  57e8cbec
- Z
  Support CUDA Graph on ParallelExecutor (#36250) · f9591bb1
  由 Zeng Jinle 提交于 10月 08, 2021
```
* support CUDA Graph on PE

* add ut, fix CI compile

* reduce memory consumption

* fix CUDA 10 CI

* improve coverage

* improve python coverage
```
  f9591bb1
- Q
  [NPU] BatchNorm support layout of NCL and NLC, test=develop (#35668) · 7cb19f57
  由 Qi Li 提交于 10月 08, 2021
```
* [NPU] support NCL and NCL for BatchNorm, test=develop

* [NPU] remove debug files, test=develop

* update, test=develop
```
  7cb19f57
- H
  add python interface of sub_graph (#36120) · a29ff4c7
  由 huangxu96 提交于 10月 08, 2021
```
Add python interface of subgraph: 1. all_sub_graphs() 2. get_sub_graph(idx)
```
  a29ff4c7
- A
  Added oneDNN BF16 relu (#36265) · 1bd9cfef
  由 arlesniak 提交于 10月 08, 2021
```
* Added oneDNN BF16 relu

* fixed typo

* refactored test, review fixes
```
  1bd9cfef
- Z
  
  fix cast cuda implementation (#36266) · 9814f895
  由 Zeng Jinle 提交于 10月 08, 2021
  
  9814f895
07 10月, 2021 1 次提交

[OneDNN] Conv op refactor. (#36252) · e9288340

由 Adam Osewski 提交于 10月 07, 2021

* Remove unused header.

* Use ConvMKLDNNHandlerT for conv2d INT8.

* Use absolute module path to import.

e9288340

05 10月, 2021 1 次提交

Added concat BF16/FP32 BWD OneDNN kernel (#35889) · dc4d5719

由 jakpiase 提交于 10月 05, 2021

* tmp

* added concat BF16/FP32 BWD oneDNN kernel

* minor change

* minor change

* fix for CI

* added formatting

* Reverted deleting static keyword

* added reviewers suggestions

* reverted deleting concat bf16 test file

* fixed concat tests

dc4d5719

30 9月, 2021 3 次提交
- Y
  
  add slotrecord datafeed (#36099) · 0a3dbe8a
  由 yaoxuefeng 提交于 9月 30, 2021
  
  0a3dbe8a
- W
  
  fix yolo (#36240) · c12176e8
  由 wenbin 提交于 9月 30, 2021
  
  c12176e8
- A
  [NPU] modify transpose2 and index_select_grad kernels for model xlnet (#36214) · a66b9fba
  由 Aganlengzi 提交于 9月 30, 2021
```
* [NPU] modify transpose2 and index_select_grad kernels for model xlnet

* add transpose2 int64_t unit test

* add more transpose2 unit tests

* update test_transpose_op_npu.py
```
  a66b9fba
29 9月, 2021 14 次提交
- Z
  Add basic support for CUDA Graph (#36190) · 21b93c3d
  由 Zeng Jinle 提交于 9月 29, 2021
```
* add basic support for CUDA Graph

* fix ci compile error

* fix LOG print, fix windows CI

* follow comments and update

* small fix for default ctor

* fix rocm compile error

* fix CPU compile error
```
  21b93c3d
- L
  fix cusparse compile problem, test=develop (#36199) · 3eb50715
  由 Liu-xiandong 提交于 9月 29, 2021
```
* fix cusparse compile problem, test=develop

* Modify file permissions
```
  3eb50715
- L
  Spinlock (#36030) · a9ea41c5
  由 liutiexing 提交于 9月 29, 2021
```
* add align for WorkQueue

* add spinlock

* merge spinlock
```
  a9ea41c5
- Y
  
  add slot record dataset (#36200) · 79bd5f90
  由 yaoxuefeng 提交于 9月 29, 2021
  
  79bd5f90
- Z
  [npu] add box coder (#36171) · 83578cfa
  由 zhulei 提交于 9月 29, 2021
```
* [npu] add box coder

* [npu] add box coder
```
  83578cfa
- P
  
  fix bug of top_k npu op (#36175) · 2b8fd704
  由 pangyoki 提交于 9月 29, 2021
  
  2b8fd704
- Z
  [NPU] Add group norm (#35937) · c79de728
  由 zhulei 提交于 9月 29, 2021
```
* [NPU] Add group norm

* [NPU] Add group norm

* [NPU] Add group norm

* [NPU] Add group norm

* [NPU] Add group_norm op
```
  c79de728
- A
  [NPU] mod for model bert (#36165) · 7bddf2e8
  由 Aganlengzi 提交于 9月 29, 2021
```
* merge conflict of paddle_gtest_main.cc

* modify FLAGS_npu_precision_mode and default not to call aclSetCompileopt
```
  7bddf2e8
- Y
  
  Implement the grad and enhance the cache of norm_convolution fusion ops. (#36168) · 767050d9
  由 Yiqun Liu 提交于 9月 29, 2021
  
  767050d9
- Z
  
  remove wait if no fetch (#36150) · b3d2dc7b
  由 Zeng Jinle 提交于 9月 29, 2021
  
  b3d2dc7b
- B
  
  fix nullptr block in op_teller (#36197) · 667bf188
  由 baoachun 提交于 9月 29, 2021
  
  667bf188
- Z
  
  refine case when thread_num = 1 (#36201) · 7e60cc63
  由 Zeng Jinle 提交于 9月 29, 2021
  
  7e60cc63
- L
  
  Add fused_dropout wrapper to ease use. (#36185) · 092d45c3
  由 Li Min 提交于 9月 29, 2021
  
  092d45c3
- R
  
  [ROCM] bugfix for bilinear_interp_v2_grad (#36160) · 5e1d0b5c
  由 ronnywang 提交于 9月 29, 2021
  
  5e1d0b5c
28 9月, 2021 11 次提交

L
Add sparse_attention api, test=develop (#35676) · 6b587e93
由 Liu-xiandong 提交于 9月 28, 2021
```
Add sparse_attention OPs, python api will be added in next pr
```
6b587e93

add API paddle.linalg.eig (#35674) · bc7e2b92

由 Lijunhui 提交于 9月 28, 2021

* Add paddle.linalg.eig op

* remove comments

* remove comments

* extend batch_size to the origin

* add real times complex functor & destroy the backward complex output bug

* terminate output diff when input real tensors

* correct tiny doc errors

* move functions from eig_helper to svd_helper and remove eig_helper

* remove tensor.Resize

* remove no longer used code

* use existing lapack functions

* reply review comments 21/27

* remove .cu as this op is only executed on CPU

* remove const_cast & add const in argument list for read-only references

* fix sample code error in CI

* remove template typename Tbase and more

* remove eig exposure in paddle.*

* add 'name=None' in eig python implementation

* handle the unittest

* try to solve the unittest

* solve CI coverage

* remove no longer used code

* polish API doc and more

* reply review comments

* polish unittest, commit plan B

* polish unittest

bc7e2b92

R

[ROCM] bugfix for arg_min_max (#36098) · 36791fdd
由 ronnywang 提交于 9月 28, 2021

36791fdd
T
[HeterPs]ps gpu dump (#36157) · 97d30602
由 Thunderbrook 提交于 9月 28, 2021
```
* ps gpu dump

* remove log
```
97d30602

[hybrid] seed and dropout op support force-cpu (#35820) · 58c8f6b3

由 xiayanming 提交于 9月 28, 2021

* [HIP] fix op not support AMD GPU bug, the flag PADDLE_WITH_ROCM is invalid

* [HIP] fix op not support AMD GPU bug, the flag PADDLE_WITH_ROCM is invalid

* [HIP] fix op not support AMD GPU bug

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] seed and dropout op support force-cpu

* [hybrid] fix seed ci failed issue

* add AsExtra for force_cpu of seed op

58c8f6b3

【Bug fix】Fix dygraph double grad dtype error (#36125) · af4f018a

由 Jiabin Yang 提交于 9月 28, 2021

* fix dygraph double grad dtype error when calling for high differential senario

* reinvoke ci

* add test for partial_engine.cc

af4f018a

L

reduce calls to SizeOfType (#36110) · c719add7
由 Leo Chen 提交于 9月 28, 2021

c719add7
G

fix bug of reduce_sum when src_dtype != dst_dtype and reduce_num == 1 (#36123) · d5268a6e
由 Guoxia Wang 提交于 9月 28, 2021

d5268a6e
Z

rename scale loss grad (#36162) · ad128144
由 Zeng Jinle 提交于 9月 28, 2021

ad128144

Add paddle.device.cuda.get_device_properties (#35661) · 4cbed9e5

由 Yanxing Shi 提交于 9月 28, 2021

* Initial Commit

* add unittest and add error information

* modify doc

* fix some error

* fix some word

* fix bug cudaDeviceProp* and modify error explanation

* fix cudaDeviceProp* error and unnitest samples

* fix hip error and PADDLE_WITH_HIP

* update style

* fix error is_compiled_with_cuda

* fix paddle.device.cuda.get_device_properties

* fix error for multi thread safe

* update style

* merge conflict

* modify after mentor review

* update style

* delete word

* fix unittest error for windows

* support string input and modify some code

* modify doc to support string input

* fix error for express information

* fix error for express information

* fix unnitest for windows

* fix device.startswith('gpu:')

* format error and doc

* fix after review

* format code

* fix error for doc compile

* fix error for doc compile

* fix error for doc compile

* fix error for doc compile

* fix error for doc compile

* fix py2 error

* fix wrong words and doc

* fix _gpuDeviceProperties

4cbed9e5

Add Basic CINN Runner Class (#35978) · 6f18b041

由 Huihuang Zheng 提交于 9月 28, 2021

* Add Basic CINN Runner Class

* Add CinnCacheKey

* Add Cache logic and improve CinnCacheKey


* Modify as reviewer commented

* Implement hash_combine to fix MAC build.

6f18b041

Crayon鑫 / Paddle 与 Fork 源项目一致

Crayon鑫 / Paddle
与 Fork 源项目一致