提交 · a8879215aa58a5c93c86ac78ac247b6d50bf31c1 · PaddlePaddle / Paddle

27 1月, 2022 11 次提交

[PluggableDevice] Add custom kernel support based on pten kernel management (#38848) · a8879215

由 Aganlengzi 提交于 1月 27, 2022

* [Demo] custom kernel based on pten kernel

* merge and npu custom work well

* del comments

* delete other code

* fix CUDAContext

* fix not found small_vector.h

* support NPU

* fix NPUContext

* fix DeviceContext support

* add UT

* fix call

* add UT

* fix

* fix for comments and ut

* add MACRO control

* fix multi input output

* support env CUSTOM_DEVICE_ROOT

* deal with special cases

* fix for Windows

* try coverage with test_custom_kernel_dot.py

* fix test_custom_kernel_dot

* fix test_custom_kernel_dot

* fix merge

* fix merge

* fix CI

* update

* merge and fix

* remove WITH_CUSTOM_KERNEL

* fix merge

* merge and fix

* fix ut

* fix ut for mac

* add more UT

* add more UT

* fix

a8879215

fix UT test_lr_scheduler random fail (#39254) · 7e6a2190
由 zhouweiwei2014 提交于 1月 27, 2022

7e6a2190

Update passes in quant2_int8_mkldnn_pass (#38912) · 0e235e58

由 joanna.wozna.intel 提交于 1月 27, 2022

* Upadate pass in quant2_int8_mkldnn_pass

* Back to the previous scale_matmul order

* Change place of cpu_quantize_placement_pass

0e235e58

W
fix shuffle_channel_detect_pass (#39242) · af9ddeb7
由 wenbin 提交于 1月 27, 2022
```
* shuffle channel pass

* add ut

* timeout fix

* makefile fix
```
af9ddeb7
C
【Auto Parallel】Update Planner (#39201) · f2226441
由 caozhou 提交于 1月 27, 2022
```
* update planner

* update unitest

* update dist matmul

* update auto converter
```
f2226441

optimize kunlun/xpu softmax_with_cross_entropy add add unitest (#39180) · 2b9bb8bb

由 QingshuChen 提交于 1月 27, 2022

* optimize kunlun/xpu softmax_with_cross_entropy add add unitest
*test=kunlun

* minor
*test=kunlun

* minor
*test=kunlun

* minor
*test=kunlun

* minor
*test=kunlun

2b9bb8bb

Add SparseCooTensor and SparseCsrTensor (#38906) · a7edb3f3

由 zhangkaihuo 提交于 1月 27, 2022

* fix bug:
1. atten: set the default value of attn_dropout_rate to None
2. ffn: add activation parameter

* for pure fp16

* Add a SparseCsrTensor

* remove unused functional

* remove const

* remove SetMemoberTensor

* remove non_zero_nums_, the number of non zero elements of each batch can be obtained from the crows

* SparseCooTensor

* add SetMember

* merge upstream; add SetMember

* merge upstream

* merge upstream; add newline at end of file

* add newline at end of file

* remove newline at end of file

* remove newline at end of file

* stash

* user pten::framework::make_ddim

* user pten::framework::make_ddim

* merge upstream; use the latest mutable_data

* merge upstream; use the latest mutable_data

* return mutable dense tensor

a7edb3f3

C
【Auto Parallel】update dist param grad for pass (#38941) · cac6f408
由 caozhou 提交于 1月 27, 2022
```
* update dist param grad for pass

* update unitest

* update unitests

* fix conflict
```
cac6f408

[Paddle-Inference]: fix concat slice (#39096) · f080e8d5

由 Wangzheee 提交于 1月 27, 2022

* Paddle-Inference:fix_concat_slice

* Paddle-Inference:fix_concat_slice

* Paddle-Inference:fix_concat_slice

* Paddle-Inference:fix_concat_slice

* [Paddle-Inference]: fix concat slice

* [Paddle-Inference]: fix concat slice

* [Paddle-Inference]: fix concat slice

f080e8d5

H
Take/Put_along_axis more input size support (#39072) · 41a64351
由 huangxu96 提交于 1月 27, 2022
```
Support the cases that the indices shape size is larger than the arr shape size
```
41a64351
Z
[Optimizer] Add master weight for opt state_dict (#39121) · 3e6950d5
由 zhangbo9674 提交于 1月 27, 2022
```
* add master weight for opt state_dict

* check empty of master weight

* strict gpu test

* refine unittest
```
3e6950d5

26 1月, 2022 11 次提交

Add FuseBatchNormAddActPass and unittest. (#39178) · 801159ce

由 hlygit66666 提交于 1月 26, 2022

* add fuse_relu_depthwise_conv_pass unittest

* fix atol and rtol

* fix according to review

* add FuseBatchNormAddActPass and unittest

* Update test_dist_fuse_bn_add_act_pass.py

* solve conflict

801159ce

[Eager] Support imperative selected_rows_to_lod_tensor and the opposite case (#39223) · 787980b1

由 Weilong Wu 提交于 1月 26, 2022

* Added selected_rows and rw_lock to pten

* Renamed the unit test target to fix CI

* Removed Class SelectedRows in Fluid, changed include/cmake relationship, use pten::SelectedRows in Fluid

* Remove rw_lock.h,rw_lock_test.cc in fluid

* Use pten::RWLock and pten::AutoRDLock, fix CI

* Use pten::SelectedRows

* Use pten::SelectedRows

* Fix to pass NPU CI

* Selected_Rows inherits from TensorBase

* Use pten::SelectedRows, to pass NPU CI

* To fix NPU CI

* To fix NPU CI again

* Use paddle/pten/core/enforce and polish code

* Support imperative selected_rows_to_lod_tensor

* Polish code

787980b1

Q
[MLU]Add conv2d op (#39110) · 71634a61
由 qipengh 提交于 1月 26, 2022
```
* [MLU]Add conv2d op

* [MLU]fix comment

* [MLU]adapt NCHW of conv2d op
```
71634a61
Y

update uts p2 (#39232) · 83d0d853
由 yaozhixin 提交于 1月 26, 2022

83d0d853

[IPU] sync misc changes 02 (#39189) · 5df78366

由 Allen Guo 提交于 1月 26, 2022

* sync misc changes

* apply comments 01

* fix compile error

* remove is_ipu_place check

* add authors
Co-authored-by: NXiaobing Wang <xiaobingw@graphcore.ai>
Co-authored-by: NAllen Guo <alleng@graphcore.ai>
Co-authored-by: NZhixin Yao <zhixiny@graphcore.ai>
Co-authored-by: NHaicheng Jiang <haichengj@graphcore.ai>
Co-authored-by: NHan Zhao <hanzhao@graphcore.ai>

* sync changes

* restore cmake

* update ir cmake and setup.py

* update inference_lib cmake

* restore for split PR
Co-authored-by: NXiaobing Wang <xiaobingw@graphcore.ai>
Co-authored-by: NZhixin Yao <zhixiny@graphcore.ai>
Co-authored-by: NHaicheng Jiang <haichengj@graphcore.ai>
Co-authored-by: NHan Zhao <hanzhao@graphcore.ai>

5df78366

Y

update uts p3 (#39214) · eb45bb4e
由 yaozhixin 提交于 1月 26, 2022

eb45bb4e
Z

change output of backward_api (#39229) · 33b3e28a
由 zyfncg 提交于 1月 26, 2022

33b3e28a
L
Optimize layer norm forward when cols is 1024. (#39167) · 01d04be6
由 Li Min 提交于 1月 26, 2022
```
* Optimize layer_norm fwd when cols is 1024.
```
01d04be6
Y

update uts p1 (#39210) · 6efb9f59
由 yaozhixin 提交于 1月 26, 2022

6efb9f59

add sigmoid cross entropy with logits to kl2 (#38915) · fd44de58

由 houj04 提交于 1月 26, 2022

* add sigmoid cross entropy with logits to kl2. test=kunlun

* add sigmoid cross entropy with logits to kl2. test=kunlun

* follow comments. test=kunlun

fd44de58

J

sum op (#39165) · 55d6b87c
由 joeqiao12 提交于 1月 26, 2022

55d6b87c

25 1月, 2022 17 次提交
- Y
  
  change infermeta and remove makePtenTenosr in reshape (#39186) · 7613129e
  由 YuanRisheng 提交于 1月 25, 2022
  
  7613129e
- H
  Add FuseBatchNormActPass and unittest. (#39176) · 09104d02
  由 hlygit66666 提交于 1月 25, 2022
```
* add fuse_relu_depthwise_conv_pass unittest

* fix atol and rtol

* fix according to review

* Add fuse_bn_act_pass unittest

* rm others

* add fuse_bn_act_pass
```
  09104d02
- Z
  [inference] update trt convert reduce op&ut,test=develop (#39088) · 80753755
  由 Zhang Jun 提交于 1月 25, 2022
```
* [inference] update convert reduce op&ut,test=develop

* update

* update

* update

* add int32 support

* add int32 support

* add comments

* trt < 7.0 do not support int32

* test=develop

* update

* test=develop
```
  80753755
- J
  [MLU]add mlu kernel for fill_constant op (#39069) · 6e871dbc
  由 joeqiao12 提交于 1月 25, 2022
```
* [MLU]add mlu kernel for fill_constant op

* delete device_context DEPS
```
  6e871dbc
- 石
  
  fix custom ops, test=develop (#39153) · 712ccfbf
  由石晓伟提交于 1月 25, 2022
  
  712ccfbf
- F
  
  fix:the axis must be 1(channel), when the dims of bias is 1 (#39052) · f07b8cbe
  由 feng_shuai 提交于 1月 25, 2022
  
  f07b8cbe
- S
  [Custom Ops]Assert _compile_dir/includes.txt existence (#39183) · 1e515aa8
  由 sneaxiy 提交于 1月 25, 2022
```
* assert _compile_dir include file existence

* polish
```
  1e515aa8
- F
  
  [MLU]add mlu batch_norm kernel pytest (#39071) · 55164761
  由 fwenguang 提交于 1月 25, 2022
  
  55164761
- J
  [MLU]add mlu kernel for split and concat (#39020) · ac3dc0bb
  由 joeqiao12 提交于 1月 25, 2022
```
* [MLU]add mlu kernel for concat and split op

* delete device_context DEPS
```
  ac3dc0bb
- Y
  
  [fleet_executor] Dist model run method Implementation (#39194) · 20e23e1b
  由 Yuang Liu 提交于 1月 25, 2022
  
  20e23e1b
- B
  
  fix_stage3_fp16 (#39171) · 8bb509d5
  由 Baibaifan 提交于 1月 25, 2022
  
  8bb509d5
- H
  [Dygraph] Support param groups in grad_clip (#39175) · b0cca48e
  由 Haohongxiang 提交于 1月 25, 2022
```
* support param groups in grad_clip

* update

* modify for review
```
  b0cca48e
- K
  
  restore gloo (#39163) · faf517b2
  由 kuizhiqing 提交于 1月 25, 2022
  
  faf517b2
- T
  
  fix test_refactor_op_xpu, *test=kunlun (#39168) · 55418d3f
  由 TTerror 提交于 1月 25, 2022
  
  55418d3f
- N
  
  [pnorm] fix bug in fp16 & optimize memory (#39011) · 3825b40f
  由 Noel 提交于 1月 25, 2022
  
  3825b40f
- C
  【Auto Parallel】Update reshard for complete (#39073) · 529f1425
  由 caozhou 提交于 1月 25, 2022
```
* update reshard for newest completion

* update unitest

* merge newest
```
  529f1425
- Z
  
  Fixed CI failure with test_eager_trace_op (#39164) · 1efeec1d
  由 Zhanlue Yang 提交于 1月 25, 2022
  
  1efeec1d
24 1月, 2022 1 次提交

[autograd] static Jacobian pass tests. (#39007) · d43655ba

由 Tongxin Bai 提交于 1月 24, 2022

* [autograd] static Jacobian pass tests.

* [autograd] apply CR suggested changes.

* [autograd] more tests.

* [autograd] add CPUPlace in tests.

* [autograd] bug fixes.

* [autograd] reformatted.

d43655ba

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功