提交 · c3f3643b26a5bf62e4dfea0d694c15d0cb397af9 · 月光在发光 / Paddle

03 3月, 2022 19 次提交

W
EmbEltwiseLayernorm fix (#40015) · c3f3643b
由 wenbin 提交于 3月 03, 2022
```
* emb fix

* fix trt6 compile

* fix half

* absolute error fix
```
c3f3643b

Modified sigmoid by the elementwise interface. (#39898) · 5d9e11a4

由 huangxu96 提交于 3月 03, 2022

* Modified sigmoid by elementwise interface.

* using TensorReduceImpl to repalce Sum function

* using reduceimpl to calculate the norm variable

* Removed useless code

5d9e11a4

Add support of int16 for gather op. (#40052) · 3e56e816

由 Li Min 提交于 3月 03, 2022

* add support of int16 for gather op.

* Recover formats.

* Recover formats.

* fix.

* Fix format.

* Fix format.

3e56e816

X
[phi] transfer pad kernel into phi and pass the test_pad_op (#40012) · 9f74b84e
由 xiongkun 提交于 3月 03, 2022
```
* add pad forward

* fix error

* transfer pad and pass the test_pad_op
```
9f74b84e
L

add communication api for ProcessGroupNCCL (#40097) · b565b349
由 lilong12 提交于 3月 03, 2022

b565b349
C

fix output var may be nullptr and cause segment fault bug (#40079) · 2ffa6436
由 chentianyu03 提交于 3月 03, 2022

2ffa6436

[PHI] Code auto-generate for Sparse API (#40060) · 31d3d857

由 zyfncg 提交于 3月 03, 2022

* suppport sparse api in yaml

* support auto-gen code of sparse api

* do some refactor

* add unittest test_sparse_conv_api

* add unitest file
Co-authored-by: Nzkh2016 <zhangkaihuo@baidu.com>

31d3d857

Workqueue threadnames (#40035) · b8a16911

由 liutiexing 提交于 3月 03, 2022

* add align for WorkQueue

* add spinlock

* merge develop

* merge

* Add EventsWaiter

* Revert "Add EventsWaiter"

This reverts commit e206173aa9be7401b83a53581627bfaf557c8fb2.

* Set thread name for WorkQueue

* Add thread names

* fix ut
Co-authored-by: Nliutiexing <liutiexing@google.com>

b8a16911

C

move gather_tree infer shape (#40082) · 3779e807
由 crystal 提交于 3月 03, 2022

3779e807
F
[Phi] move gaussian_random (#39932) · 00bbb8c5
由 furnace 提交于 3月 03, 2022
```
[Phi] move gaussian_random kernel
```
00bbb8c5
B

change_ASP_sharding_option (#40028) · 815f7a67
由 Baibaifan 提交于 3月 03, 2022

815f7a67
Z

bugfix in is_xpu_support_op (#40070) · 34d93bee
由 zhangxiaoci 提交于 3月 03, 2022

34d93bee

Support slim eager (#39874) · da47544c

由 Jiabin Yang 提交于 3月 03, 2022

* eager, test=develop

* fix bug, test=develop

* eager, test=develop

* merge legacy to fluid

* eager, test=develop

* eager, test=develop

* Refactor TensorAdd func by template and remove gradient_accumulation in eager

* Remove needless target name

* eager, test=develop

* eager, test=develop

* Use overload instead of template

* Remove legacy code

* Remove legacy code

* selectedrows, test=develop

* Remove DataType test

* eager, test=develop

* eager, test=develop

* support gan, test=develop

* Using Tensor directly instead of using EagerTensor

* support gradient_accumulation

* make test_imperative_lod_tensor_to_selected_rows longer

* make test_imperative_lod_tensor_to_selected_rows longer

* refine code

* ptb, test=develop

* Rename all EagerTensor to Tensor

* Rename some EagerTensor to Tensor

* rename EagerTensor to EagerVariable

* eager, test=develop

* eager, test=develop

* eager, test=develop

* eager, test=develop

* add more test

* eager, test=develop

* Support copiable selected rows and merge develop

* save load, eager, test=develop

* save load, eager, test=develop

* refine, test=develop

* remove useless _set_value method

* refine, test=develop

* refine, test=develop

* revert static_runner, test=develop

* EagerTensor to Tensor, test=develop

* refine, test=develop

* refine, test=develop

* clear grad, test=develop

* merge, develop

* merge, develop

* merge, test=develop

* merge, test=develop

* Support quant and part of slice

* support legacy static save

* extend slim tests time

* remove imperative on inference

* remove imperative on inference

* merge develop

* fix typo

* fix typo

* split slice related code into 2 part for imperative and eager

* split slice from inference

* split slice from inference

* fix test_tensor_register_hook
Co-authored-by: NWang Huan <wanghuan29@baidu.com>
Co-authored-by: NWeilong Wu <veyron_wu@163.com>
Co-authored-by: Nwanghuancoder <wanghuancoder@163.com>

da47544c

Z

adjust the args checking of backward in yaml (#40091) · d9884e20
由 zyfncg 提交于 3月 03, 2022

d9884e20
N
Modified Reduce for XPU2 (#38918) · 909d1e61
由 niuliling123 提交于 3月 03, 2022
```
1. set xpu2 block_size = 64
2. fix a bug when reduce_num is too large
```
909d1e61
Z
Implement SparseConv3d kernel (#39784) · 6bf85eaf
由 zhangkaihuo 提交于 3月 03, 2022
```
* sparse conv3d: gpu code
```
6bf85eaf
Z

[Eager][YAML] Supported array-type parsing for output tensors (#40058) · 71c69507
由 Zhanlue Yang 提交于 3月 03, 2022

71c69507

Move bn to pten (#39347) · ebd0f512

由 hong 提交于 3月 03, 2022

* add bn cpu version; test=develop

* move batch norm to pten

* move batch norm to pten; test=develop

* fix bug; test=develop

* fix func::tranpose depend bug; test=develop

* fix compile bugs; test=develop

* fix use_op batch_norm bug; test=develop

* fix cudnn bn add relu test; test=develop

* fix pten context build and double grad bug; test= develop

* remve useless code; test=develop

* add batch norm gpu fp16 support; test=develop

* fix test bn op bug; test=develop

* remove output dtype set; test=develop

* fix bug; test=develop

* fix bug; test=develop

* fix applay pass to program bug; test=develop

* revert to develop; test=develop

* fix rocm bug; test=develop

* revert operator to develop; test=develop

* fix pre_commit; test=develop

* fix statci check error; test=develop

* resolve conflict; test=develop

* ana batch norm bug;

* revert batch norm op

* resolve conlict

* fix nan inf and speed bug; test=develop

* fix bug; test=develop

* fix error; test=develop

* test expand op; test=develop

* fix bug; test=develop

* resolve confilct

* resolve confilct; test=develop

* polish code; test=develop

* polish code; test=develop

* change mutable data to ctx alloc; test=develop

* make format same with ci; test=develop

* fix format error with ci; test=develop

ebd0f512

L
Add the implementation of Gloo for ProcessGroup (#39892) · c16f85f9
由 lilong12 提交于 3月 03, 2022
```
* add pg_gloo
```
c16f85f9

02 3月, 2022 21 次提交

L
Replacing dropout eval eigen usage by cuda kernel (#40053) · 272b32fd
由 Li Min 提交于 3月 02, 2022
```
* Replacing dropout eval eigen usage by cuda kernel
```
272b32fd
F
[MLU] add mlu ci script (#39805) · a8e02ef1
由 fwenguang 提交于 3月 02, 2022
```
* [MLU] add mlu ci script

* Update CMakeLists.txt
```
a8e02ef1

Move sgd to phi (#40045) · f3d54e2e

由 hong 提交于 3月 02, 2022

* move sgd to phi; test=develop

* update

* add sgd kernel; test=develop

f3d54e2e

Z
Adjust GPU Arches for next level Whl release strategy (#39910) · 3fc698fb
由 Zhanlue Yang 提交于 3月 02, 2022
```
* Adjust GPU Arches for Whl releases

* Adjusted CUDA arches

* fixed minor issue

* adjusted gpu arches
```
3fc698fb
W
modify infershape of yolo_box (#40056) · ebc6959c
由 wangxinxin08 提交于 3月 02, 2022
```
* modify infershape of yolo_box
```
ebc6959c
A
[IPU] update dockerfile (#40061) · 7ef61789
由 Allen Guo 提交于 3月 02, 2022
```
* update dockerfile for ipu

* update comments, test=document_fix
```
7ef61789
L
add check for backward hook (#40041) · 1980e33a
由 Leo Chen 提交于 3月 02, 2022
```
* add check for backward hook

* refine ut
```
1980e33a
S
Move gather.h/gather.cu.h/scatter.h/scatter.cu.h to the phi library (#40043) · 09258040
由 sneaxiy 提交于 3月 02, 2022
```
* move gather.h gather.cu.h scatter.h scatter.cu.h to phi library

* fix CI

* fix rocm ci
```
09258040
S

vec scale kernel (#40011) · 2e6548a9
由 sneaxiy 提交于 3月 02, 2022

2e6548a9
Y
[Phi]Move elementwise function to funcs directory (#39986) · 5898e9ab
由 YuanRisheng 提交于 3月 02, 2022
```
* move elementwise function to funcs directory

* fix compile bugs

* modify according to comment
```
5898e9ab
A
[XPU] Fix Phi Kernel cache problem in operator.cc (#40044) · 66196573
由 Aurelius84 提交于 3月 02, 2022
```
* [XPU] Fix Phi Kernel cache problem in operator.cc

* fix typo
```
66196573

Move transpose to pten (#39327) · 7a857924

由 hong 提交于 3月 02, 2022

* immigrate_transpose_to_pten cpu kernel only; test=develop

* fix bug; test=develop

* add transpose cuda api

* bug fix;

* fix bugs

* fix bugs; test=develop

* bug fix;

* move transepose to pten; test=develop

* fix bug; test=develop

* fix bugs; test=develop

* add transpose grad fp16 support; test=develop

* fix bug; test=develop

* fix npu bug; test=develop

* fix nemul = 0 bug; test=develop

* add fp16 support; test=develop

* fix data type register bug; test=develop

* fix transpose bug; test=develop

* update transpose

* fix transpose bug; test=develop

* remove useless code; test=develop

* remove useless code; test=develop

* fix transpose alias bug; test=develop

* polish code; test=develop

* resolve confict; test=develop

* resolve confilct; test=develop

* recover prepared operator; test=develop

* fix bug; test=develop

* polish code; test=develop

* fix bug; test=develop

* fix bug; test=develop

7a857924

Move BroadcastTensors OP to phi (#40047) · 2a5590a1

由 From00 提交于 3月 02, 2022

* Move BroadcastTensors OP to phi

* Remove mutable_data in impl

* Move BilinearTensorProductInferMeta to multiary.h/cc

2a5590a1

Z
The backward code of Sparse Conv3d (#40054) · 8492d3bb
由 zhangkaihuo 提交于 3月 02, 2022
```
Sparse Conv3d backward code
```
8492d3bb
L

run recompute's real backward with amp disabled (#40042) · 28795771
由 Leo Chen 提交于 3月 02, 2022

28795771

new fleet_desc builder (#39948) · 1c4e3e5d

由 ziyoujiyi 提交于 3月 02, 2022

* delete gloo connect retry

* the_one_ps dirs reconstruct

* .

* .

* create the_one_ps dirs

* create the_one_ps dirs

* create the_one_ps dirs

* create the_one_ps dirs

* create the_one_ps dirs

* create the_one_ps dirs

* the one ps dirs modify

* the one ps dirs modify

* the one ps dirs modify

* the one ps dirs modify

* refactor ps optimize

* refactor ps optimize

* refactor ps optimize

* .

* .

* .

* .

* .

* .

* refactor theoneps

* the_one_ps

* add ps pass unittest

* add ps pass unittest

* ps unitest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* ps unittest frame

* add cpu_async_ps_mode test

* add cpu_async_ps_mode test

* add cpu_async_ps_mode test

* ps unittest ready

* ps unittest ready

* solve dist_pass init conflict

* solve import CommContext error

* unittest ok

* implement AllocateFrom

* solve setup.py.in conflict

* solve conflict

* solve conflict

* solve conflict

* .

* .

* cpu-async-ps minimize test ok & gpu minimize test ok

* add heter 2stage unittest

* add heter 2stage unittest

* add heter 2stage unittest

* sync/geo test ok & fix heter_worker program ok

* .

* new fleet desc generator

* new fleet_desc builder

* new fleet_desc builder

* .

* .

* correct ps.proto compile

* .
Co-authored-by: Nzkh2016 <zhangkaihuo@baidu.com>

1c4e3e5d

P
support checking `phi` directory in CI op benchmark (#40026) · f30b3f81
由 pangyoki 提交于 3月 02, 2022
```
* support phi checking in CI op benchmark

* add sparse/gpu

* remove h file in cpu directory
```
f30b3f81
H

[Infrt]add phi kernel dialect (#39726) · 07dad6d6
由 huzhiqiang 提交于 3月 02, 2022

07dad6d6
Z
[bf16] add bf16 kernel: softmax & log_softmax (#39999) · 4a4215ff
由 zhangbo9674 提交于 3月 02, 2022
```
* add softmax log_softmax

* refine rocm

* refine unittest
```
4a4215ff
J
[Auto Parallel] Adapt Partitioner & DistOp for ERNIE3.0 Inference and cache (#39895) · c9cd47d9
由 JZ-LIANG 提交于 3月 02, 2022
```
* adapot dist op

* add dist_fill_constant_batch_size_like

* remvoe print

* update compitable

* add unitest
```
c9cd47d9
C
【phi】migrate gather_tree,reduce_prod to phi (#39844) · 6af2729e
由 crystal 提交于 3月 02, 2022
```
* move to phi

* migrate gather_tree_op into phi

* move reduce_prod tp phi

* optimize code
```
6af2729e

月光在发光 / Paddle 与 Fork 源项目一致

月光在发光 / Paddle
与 Fork 源项目一致