提交 · 14658d8f659a6c03ce9c5b1bdeb01cf69d7f800a · PaddlePaddle / Paddle

29 12月, 2021 10 次提交
- H
  Fix buddy allocator random CI failure (#38545) · 14658d8f
  由 Huihuang Zheng 提交于 12月 29, 2021
```
Fix Buddy Allocator random CI failure due to machine environment.
```
  14658d8f
- 王
  
  [infrt] fix infrt ci test. test=develop, test=infrt (#38533) · 9d5b665c
  由王明冬提交于 12月 29, 2021
  
  9d5b665c
- Y
  
  add top k v2 operator, test=kunlun (#38434) · d22f92ad
  由 ykkk2333 提交于 12月 29, 2021
  
  d22f92ad
- S
  
  fix reduce_max/reduce_min bug (#38476) · 995332ef
  由 Shang Zhizhou 提交于 12月 29, 2021
  
  995332ef
- T
  add argsort/scatter for kunlun (#38345) · 4643baa7
  由 TTerror 提交于 12月 29, 2021
```
* add argsort/scatter for kunlun

* update test_scatter

* update xpu.cmake

* update xpu.cmake

* fix scatter
```
  4643baa7
- S
  
  fix lamb beta1pow beta2pow update (#38518) · 3672480b
  由 sneaxiy 提交于 12月 29, 2021
  
  3672480b
- T
  
  reduce compile time of amax and amin (#38534) · 72a41e50
  由 Tao Luo 提交于 12月 29, 2021
  
  72a41e50
- S
  
  add nccl func of NCCL 2.11 (#38519) · 4853ab0a
  由 sneaxiy 提交于 12月 29, 2021
  
  4853ab0a
- L
  
  code clean (#38550) · 206a8f6c
  由 limingshu 提交于 12月 29, 2021
  
  206a8f6c
- W
  
  [fleet_executor] remove SetCreatingFlag (#38539) · 9171aaa0
  由 WangXi 提交于 12月 29, 2021
  
  9171aaa0
28 12月, 2021 13 次提交

L
Support multi-output feature for elementwise (#38410) · 48f061fb
由 limingshu 提交于 12月 28, 2021
```
* first commit

* pass ctest of  elementwise_div_grad
```
48f061fb

Utilize StreamSafeCUDAAllocator to support fast GC in new executor (#37642) · 0c7153a4

由 From00 提交于 12月 28, 2021

* fix reshape move storage error

* remove needless set type

* alloc tensor by shared storage

* Utilize StreamSafeCUDAAllocator to support fast GC in new executor

* Fix compile error for Windows and ROCm

* Fix compile error for Windows

* Modify UT stream_safe_cuda_alloc_test

* Modify UT stream_safe_cuda_alloc_test

* Rewrite fast GC

* Rewrite fast GC

* Fix compile error for BOOST_GET_CONST

* Fix compile error for BOOST_GET_CONST

* Changes default stream for StreamSafeCUDAAllocator

* Fix a small CI error

* Remove some redundant code

* Fix conflict

* Fix compile error for ROCm

* Fix Windoes CI error

* Fix CI error

* Remove some unnecessary code

* Fix CI error

* Add UT for fast GC

* Fix CI error

* add device-agnostic stream class

* add stream.h

* fix ut

* fix cpu compile

* Use RWLock in GetAllocator

* Fix CI error
Co-authored-by: NChen Weihang <chenweihang@baidu.com>
Co-authored-by: Nzhiqiu <chenqiuliang@baidu.com>

0c7153a4

Support test basic of Var and Layer (#38426) · 1fb80a6a

由 Jiabin Yang 提交于 12月 28, 2021

* Rearranged Eager AutoCodeGen directory structure

* Removed USE_OP in Eager AutoCodeGen

* Enabled generation for Operators without Grad/Inputs/Outputs

* Resolved operators without input

* Fixed merge conflicts

* Enabled Eager AutoCodeGen for 10+ more operators

* Refactored Eager AutoCodeGen with more organized helper objects

* Enabled Eager AutoCodeGen for operators with multiple OpBases

* Adjusted Eager AutoCodeGen to Enable Passing Output Tensor as Input Argument

* Handled Dispensable Inputs/Outputs in Eager AutoCodeGen

* Adjusted function generation/call between Python-C API & Dygraph API

* Synchronized auto-generated Python-C API with Dygraph Forward Functions

* support more eager tensor api

* fix merge compile error

* fix compile error and fit develop code

* support pure CPU

* fix some logic error in eager_mode

* support _varbase_creator in eager mode

* Added safe_initialized interface to EagerTensor for use in processing dispensable inputs

* for eager mode

* refine

* support multiple constructor for eager tensor

* add place related code

* polish code

* specific randint with dtype of int64

* Support pure cpu test

* eager logic

* refine test in pure cpu

* eager logic

* eager logic

* eager logic, test=develop

* skip core.eager when in inference, test=develop

* refine, test=develop

* refine, test=develop

* call RetainGrad after run forward kernel, test=develop

* refine, test=develop

* support dygraph util, meta, guard test

* support inference test

* refine test and fix initializer failed

* support create varbase and fix retain grad error

* fix windows error

* support test code coverage

* support test code coverage

* support test code coverage
Co-authored-by: Njim19930609 <jim19930609@gmail.com>
Co-authored-by: NWang Huan <wanghuan29@baidu.com>

1fb80a6a

Z
refactor matmul directory in pten (#38227) · 982bf444
由 zyfncg 提交于 12月 28, 2021
```
* refactor matmul directory in pten

* fix merge conflict
```
982bf444

Add API and op for take_along_axis (#38396) · 3310f519

由 huangxu96 提交于 12月 28, 2021

* add API and op for take_along_axis

* fix compile dependency problem and add example code and doc

* add unitest

* delete some code for CI coverage

* fix code style problem

* fix as review

3310f519

G

fix adamw epsilon in cuda kernel (#37746) · 6f1bb3d6
由 Guoxia Wang 提交于 12月 28, 2021

6f1bb3d6
T
Add Amax and Amin API (#38417) · 340dfb26
由 Tao Luo 提交于 12月 28, 2021
```
* add amax/amin

* support axis is list
```
340dfb26

[pten] remove in_type arg in cast kernel (#38486) · 0637b9a6

由 chentianyu03 提交于 12月 28, 2021

* remove intype arg in cast kernel

* modify conj config in api.yaml by dictionary order

* rm unused code in cast_kernel.cu

0637b9a6

add reduce_prod_xpu. fix reduce_mean_xpu bug. (#38481) · 78836bb7

由 houj04 提交于 12月 28, 2021

* add reduce_prod_xpu. fix reduce_mean_xpu bug.

* iadd reduce_prod_xpu. fix reduce_mean_xpu bug. test=kunlun

78836bb7

L
[new-exec] add completion_nofifier (#38447) · 404a4a6a
由 Leo Chen 提交于 12月 28, 2021
```
* add completion_nofifier

* fix bug

* unregist event waiter
```
404a4a6a

add mul_lstm_fuse_pass ut (#37795) · 1db61c3e

由 baoachun 提交于 12月 28, 2021

* add mul_lstm_fuse_pass ut

* update mul_lstm_fuse_pass ut

* update ut

* update ut

* update ut

* add CPU ut cmake setting

* update ut

1db61c3e

Z

Fixed issue with offset,test=allcases (#38506) · dc30ad1d
由 Zhanlue Yang 提交于 12月 28, 2021

dc30ad1d
L

Add constructor for fused dropout param to ease use. (#38475) · f9e8a775
由 Li Min 提交于 12月 28, 2021

f9e8a775

27 12月, 2021 13 次提交
- W
  
  [fleet_executor] Add task loop thread pool (#38420) · dba59db7
  由 WangXi 提交于 12月 27, 2021
  
  dba59db7
- Y
  [PTen]move reshape kernel according to new directory (#38432) · 49216134
  由 YuanRisheng 提交于 12月 27, 2021
```
* move reshape

* fix compile bugs

* delete manipulation file

* fix compile bugs
```
  49216134
- P
  fix accumulator bug when multiple inplace OPs are executed continuously (#38406) · 113c8b93
  由 pangyoki 提交于 12月 27, 2021
```
* fix accumulator bug

* fix unittest
```
  113c8b93
- S
  [BugFix]Fix bug in pfp16 in DataParallel (#38378) · e8e47581
  由 ShenLiang 提交于 12月 27, 2021
```
* fix bug in pfp16

* fix hip

* fix hip
```
  e8e47581
- B
  
  update mkldnn matmul_transpose_reshape fuse pass ut (#38467) · 9cfdae91
  由 baoachun 提交于 12月 27, 2021
  
  9cfdae91
- B
  add matmulv2_transpose_reshape_pass ut (#37416) · f664a533
  由 baoachun 提交于 12月 27, 2021
```
* update mkldnn matmul_v2_transpose_reshape_fuse_pass ut

* update mkldnn matmul_v2_transpose_reshape_fuse_pass ut

* update ut

* update ut
```
  f664a533
- L
  add device-agnostic stream class (#38391) · 6b5e33b4
  由 Leo Chen 提交于 12月 27, 2021
```
* add device-agnostic stream class

* add stream.h

* fix ut

* fix cpu compile
```
  6b5e33b4
- S
  
  refine float16 implementation (#38439) · 78375990
  由 sneaxiy 提交于 12月 27, 2021
  
  78375990
- L
  Support multi-outputs feature for broadcast ops (#38329) · 89d38f55
  由 limingshu 提交于 12月 27, 2021
```
* No harm to KP

* Pass the compile stage

* change the WriteData function

* fix template bugs and pass ctest of current elementwise

* for passing partial template specialization of tempalte function in CI-ROCm

* To make 'WriteData' funtion flexible.

* a less harmful way to support multi-output

* a less harmful way to support multi-output
```
  89d38f55
- C
  
  remove npu related impl (#38428) · f1d56b77
  由 Chen Weihang 提交于 12月 26, 2021
  
  f1d56b77
- C
  [PTen] Move cast kernel impl (#38382) · 1fb734d7
  由 Chen Weihang 提交于 12月 26, 2021
```
* rename to api to copy_to

* revert needless change

* polish format
```
  1fb734d7
- B
  
  add attr check for infer in batch_norm_act mkldnn fuse pass (#38443) · 04527ee3
  由 baoachun 提交于 12月 27, 2021
  
  04527ee3
- G
  
  gelu using normcdf for cudnn (#38450) · 37022482
  由 Guoxia Wang 提交于 12月 27, 2021
  
  37022482
26 12月, 2021 4 次提交
- C
  [PTen] Move copy kernel impl (#38421) · 73819658
  由 Chen Weihang 提交于 12月 26, 2021
```
* add register general kernel marco

* move copy kernel impl

* revert needless change

* polish details

* fix xpu compil faild

* fix xpu compile failed

* polish format
```
  73819658
- Z
  
  improve forward performace (#38279) · acef85b2
  由 Zhang Ting 提交于 12月 26, 2021
  
  acef85b2
- C
  Fix renorm op include error and format error (#38451) · e6c3f64f
  由 Chen Weihang 提交于 12月 25, 2021
```
* remove needless header

* remove needless header

* adjust header order
```
  e6c3f64f
- Z
  [Unify Tensors PR #2] Replaced pten::LoD with paddle::framework::LoD (#38275) · bbe879fc
  由 Zhanlue Yang 提交于 12月 26, 2021
```
* Replaced pten::LoD with paddle::framework::LoD

* Overrided CPUVector with CUDAVector

* Refactored paddle::framework::Vector
```
  bbe879fc

PaddlePaddle / Paddle 1 年多 前同步成功

PaddlePaddle / Paddle
1 年多前同步成功