提交 · 1db188f318ae0b0292984e08afd626898e3170da · Crayon鑫 / Paddle

02 3月, 2022 10 次提交
- A
  [IPU] update ipu unittests p0 (#39707) · 1db188f3
  由 Allen Guo 提交于 3月 02, 2022
```
* update ipu UTs part0

* rename UT

* sync api changes

* update uts for new api

* use_ipumodel() as classmethod
```
  1db188f3
- J
  
  add logic kernel for mlu (#39940) · bc113e10
  由 joeqiao12 提交于 3月 02, 2022
  
  bc113e10
- Y
  [fleet_executor] Add entrance of FleetExecutor in AnalysisPredictor for... · 244ae318
  由 Yuang Liu 提交于 3月 02, 2022
```
[fleet_executor] Add entrance of FleetExecutor in AnalysisPredictor for distributed inference (#39992)
```
  244ae318
- Q
  [MLU] adapt matmul op (#39727) · b4d931e8
  由 qipengh 提交于 3月 02, 2022
```
* [MLU] adapt matmul op

* [MLU] fix phi namespace
```
  b4d931e8
- F
  
  [MLU] add transpose2 mlu kernel (#39994) · 4cab812e
  由 fwenguang 提交于 3月 02, 2022
  
  4cab812e
- B
  
  add_new_comm_primitive (#40040) · 4e00d2bb
  由 Baibaifan 提交于 3月 02, 2022
  
  4e00d2bb
- L
  
  fix unittests for eignvalsh (#39841) · aa47297a
  由 lkylkylky 提交于 3月 02, 2022
  
  aa47297a
- optimize CUDA implementaion of randint OP (#39952) · fb635089
  由 zhouweiwei2014 提交于 3月 02, 2022
```
* change CUDA implementaion of randint OP,move distribution common func to phi

* fix CI

* fix CI
```
  fb635089
- W
  [Eager] open eager when WITH_PYTHON (#39979) · 9af72957
  由 wanghuancoder 提交于 3月 02, 2022
```
* open eager when WITH_PYTHON, test=develop

* refine, test=develop

* refine, test=develop

* add DWITH_PYTHON for gen_fluid_lib, test=develop
```
  9af72957
- W
  
  [Eager] Support gnn ptb_rnn in eager mode (#39993) · dbcf8797
  由 Weilong Wu 提交于 3月 02, 2022
  
  dbcf8797
01 3月, 2022 10 次提交
- fix bug of paddle.to_tensor and paddle.moveaxis (#39662) · 4617c1b2
  由 zhouweiwei2014 提交于 3月 01, 2022
```
* fix bug of paddle.to_tensor and paddle.moveaxis

* fix CI
```
  4617c1b2
- A
  
  fix compiling and running with ipu (#39920) · 69ab2700
  由 Allen Guo 提交于 3月 01, 2022
  
  69ab2700
- J
  Add mobilenetv3_large performance test for bf16 and int8 (#39738) · eb7c211a
  由 joanna.wozna.intel 提交于 3月 01, 2022
```
* Add mobilenetv3_large performance test

* Disable the BF16 test if the device does not support BF16 computations

* Change test timeout
```
  eb7c211a
- Z
  [bf16] add bf16 kernel: layer_norm p_norm reduce_sum (#39843) · ce8ed978
  由 zhangbo9674 提交于 3月 01, 2022
```
* add layer norm

* add p norm

* add reduce sum

* refine layer norm register bf16 for cudnn811

* add bf16 cast for hip

* add unittest

* refine rocm

* refine layer_norm unittest

* refine reduce op

* refine unittest

* enhance atol for reduce unittest
```
  ce8ed978
- W
  remove conv_affine_channel_fuse_pass (#39817) · fc06be9d
  由 wenbin 提交于 3月 01, 2022
```
* remove

* pass

* more pass
```
  fc06be9d
- Z
  
  add test_warpctc_op in mac (#39983) · 25650774
  由 zhangchunle 提交于 3月 01, 2022
  
  25650774
- Z
  [bf16] add bf16 kernel: scale gather sum (#39683) · 6d26b332
  由 zhangbo9674 提交于 3月 01, 2022
```
* add scale gather sum

* refine CUDA_ATOMIC_WRAPPER ADD for bf16

* add gather unittest

* solve conflict

* add scale uinttest

* add sum unittest

* solve conflict

* refine gather unittest

* refine unittest
```
  6d26b332
- A
  [Phi] Migrate logical_and/or/not/xor into Phi (#39942) · 8c237973
  由 Aurelius84 提交于 3月 01, 2022
```
* [Phi] Migrate logical_and/or/not/xor into Phi

* fix unittest

* fix function name
```
  8c237973
- S
  [DP] Construct reducer group (#39987) · 4da841e0
  由 ShenLiang 提交于 3月 01, 2022
```
* add reducer
```
  4da841e0
- S
  Optimize the CUDA kernel in DistributedFusedLamb optimizer (#39972) · d17961ed
  由 sneaxiy 提交于 3月 01, 2022
```
* vectorize lamb kernel

* remove flags, add ut

* remove useless codes

* refine code, add param order
```
  d17961ed
28 2月, 2022 3 次提交
- Z
  PR-CI-Py3 change cpu test (#39659) · 3cb93edf
  由 zhangchunle 提交于 2月 28, 2022
```
* update;test=cpu-py3
```
  3cb93edf
- C
  [Pten->Phi PR4] Rename pten in funcs to phi (#39961) · eb42dd52
  由 Chen Weihang 提交于 2月 28, 2022
```
* rename pten_utils to phi_utils

* rename pten_utils target

* rename Pten to Phi

* replace pten with phi

* resolve conflict
```
  eb42dd52
- Z
  [bf16] Refine BF16 amp-o1 logic (#39815) · 18ee051e
  由 zhangbo9674 提交于 2月 28, 2022
```
* refine bf16 amp-o1 logic

* refine amp GLOG

* refine unittest

* refine unittest
```
  18ee051e
27 2月, 2022 1 次提交
- L
  fix pylayer problem with amp (#39950) · 282e09dc
  由 Leo Chen 提交于 2月 27, 2022
```
* fix pylayer problem with amp

* add ut

* refine code
```
  282e09dc
26 2月, 2022 1 次提交
- W
  [Eager Hook] Support GradientHook and ReduceHook, expose related interface to python (#39893) · a456dda6
  由 Weilong Wu 提交于 2月 26, 2022
```
* Support Eager Hook, expose interface to python

* Fix CI issue
```
  a456dda6
25 2月, 2022 5 次提交
- J
  
  added logsoftmax oneDNN kernel (#39793) · 584844ec
  由 jakpiase 提交于 2月 25, 2022
  
  584844ec
- Z
  
  [MLU]support launch process on mlu (#39839) · 2533cac6
  由 zn 提交于 2月 25, 2022
  
  2533cac6
- Z
  [bf16] add bf16 kernel: elementwise_add elementwise_mul elementwise_sub (#39716) · 2fedd39b
  由 zhangbo9674 提交于 2月 25, 2022
```
* add ele_add

* add ele_mul

* add ele_sub

* sovle conflict

* fix npu

* refine ele_add

* add ele_mul unittest

* refine ele_sub

* refine ci

* refine unittest
```
  2fedd39b
- J
  
  add reduce_min and reduce_max (#39899) · 44da9b42
  由 joeqiao12 提交于 2月 25, 2022
  
  44da9b42
- F
  
  [MLU] add elementwise_mul mlu kernel (#39864) · 04d324b2
  由 fwenguang 提交于 2月 25, 2022
  
  04d324b2
24 2月, 2022 8 次提交

[IPU] Update IpuStrategy Python Part (#39646) · e0409c93

由 Allen Guo 提交于 2月 24, 2022

* Update IpuStrategy Python Part

* add docs

* add add_custom_op for ipu_strategy

* fix build warning

* rm unneeded part

* clean api

* fix typo

* update option names

* update IpuStrategy doc

e0409c93

Z

[MLU]add mlu kernel for allreduce (#39788) · ce207c3a
由 zn 提交于 2月 24, 2022

ce207c3a
R

fix paddle.where torch diff (#39859) · c5ae43a2
由 ronnywang 提交于 2月 24, 2022

c5ae43a2
C
Fix unittests for eigh op (#39568) · 539fb0d7
由 crystal 提交于 2月 24, 2022
```
* fix eigh test

* modify atol and rtol
```
539fb0d7

Added nearest interp v2 BF16 FWD kernel (#39490) · 2ec943a7

由 jakpiase 提交于 2月 24, 2022

* added nearest interp v2 bf16

* disabled bilinear interp nhwc test

* added skipping UT for gpu

* added NHWC support

* removed unnecessary statements

* minor change

* CI fix

* added appropriate changes to interpolate_v1

* fix after review

* minor change

* minor change

* revert unwanted deletions

* CI fix

2ec943a7

Refactored GradNodeAccumulation data structure and behaviour (#39526) · 1abfc8dd

由 Zhanlue Yang 提交于 2月 24, 2022

* Refactored GradNodeAccumulation data structure and behaviour

* Fixed CI issues

* Fix compilation issues

* Fixed minor issues

* Reverted changes for intermediate and OverwriteOutput

* fixed minor issue

* Fixed code format issues

* Fixed CI-Coverage issue

* Fixed CI issues

1abfc8dd

H
Add Note for Place of Executor in Parallel Environment (#39063) · 867224b2
由 Huihuang Zheng 提交于 2月 24, 2022
```
Add note for Place of Executor in parallel environment
```
867224b2

[Eager] save load testcase (#39571) · 6b5749eb

由 wanghuancoder 提交于 2月 24, 2022

* eager, test=develop

* fix bug, test=develop

* eager, test=develop

* merge legacy to fluid

* eager, test=develop

* eager, test=develop

* Refactor TensorAdd func by template and remove gradient_accumulation in eager

* Remove needless target name

* eager, test=develop

* eager, test=develop

* Use overload instead of template

* Remove legacy code

* Remove legacy code

* selectedrows, test=develop

* Remove DataType test

* eager, test=develop

* eager, test=develop

* support gan, test=develop

* Using Tensor directly instead of using EagerTensor

* support gradient_accumulation

* make test_imperative_lod_tensor_to_selected_rows longer

* make test_imperative_lod_tensor_to_selected_rows longer

* refine code

* ptb, test=develop

* Rename all EagerTensor to Tensor

* Rename some EagerTensor to Tensor

* rename EagerTensor to EagerVariable

* eager, test=develop

* eager, test=develop

* eager, test=develop

* eager, test=develop

* add more test

* eager, test=develop

* Support copiable selected rows and merge develop

* save load, eager, test=develop

* save load, eager, test=develop

* refine, test=develop

* refine, test=develop

* refine, test=develop

* revert static_runner, test=develop

* EagerTensor to Tensor, test=develop

* refine, test=develop

* refine, test=develop

* clear grad, test=develop

* merge, develop

* merge, develop

* merge, test=develop

* merge, test=develop
Co-authored-by: NJiabinYang <360788950@qq.com>
Co-authored-by: NWeilong Wu <veyron_wu@163.com>

6b5749eb

23 2月, 2022 2 次提交
- J
  
  added paddle_bfloat to requirements (#39740) · 2457a7d1
  由 jakpiase 提交于 2月 23, 2022
  
  2457a7d1
- S
  Add ProcessGroupNCCL for distributed training (#39737) · 0b205817
  由 ShenLiang 提交于 2月 23, 2022
```
* add processgroup_nccl
```
  0b205817

Crayon鑫 / Paddle 与 Fork 源项目一致

Crayon鑫 / Paddle
与 Fork 源项目一致