提交 · aaabb796a8ae6cf5b4ab24997f2cb898ef7d0b48 · PaddlePaddle / Paddle

18 4月, 2022 13 次提交
- W
  [Eager] use final op in maskrcnn and hrnet (#41927) · aaabb796
  由 wanghuancoder 提交于 4月 18, 2022
```
* update

* add conv yaml

* add backward

* remove useless code

* fix bug

* fix bug

* revert fluid dygraph conv2d

* remove useless infermeta function

* fix meta fn deluplicat error

* conv using custom impl

* remove amp include

* fix bug

* use final op in maskrcnn and hrnet

* refine
Co-authored-by: Nphlrain <phliuhongyu@126.com>
```
  aaabb796
- Z
  
  add data transform config for shape and size (#41909) · 9291358d
  由 zyfncg 提交于 4月 18, 2022
  
  9291358d
- Z
  Create Tensor by paddle::empty in custom operator (#41840) · bc1c3e3e
  由 zyfncg 提交于 4月 18, 2022
```
* create tensor by empty in custom op

* fix some bug
```
  bc1c3e3e
- L
  
  [KP] Add Reduce op registry & UT for xpu_kp compilation (#41869) · b3959fe4
  由 Lijunhui 提交于 4月 18, 2022
  
  b3959fe4
- Z
  [AutoParallel] dist slice op (#41780) · 14c35a58
  由 zhaoyingli 提交于 4月 18, 2022
```
* add dist slice

* fix debug

* fix cmakelist
```
  14c35a58
- Y
  [Auto Parallel] Add links for the reference codes. (#41829) · 338e00dd
  由 Yulong Ao 提交于 4月 18, 2022
```
* [Auto Parallel] Add links for the reference codes.

* Update recorder.py

* Update trial.py

* Update test=document_fix

* Update test=document_fix
```
  338e00dd
- L
  
  fix bug for eager mode distributed training (#41841) · 34f30f79
  由 lilong12 提交于 4月 18, 2022
  
  34f30f79
- support tril_triu_grad for KL2, *test=kunlun (#41877) · 0759e99d
  由 z8hanghuan 提交于 4月 18, 2022
  
  0759e99d
- J
  [Auto parallel] Transformer MHA & FFN Fused Dist op (#41163) · ceef73c9
  由 JZ-LIANG 提交于 4月 18, 2022
```
* adapot dist op

* [Auto Parallel] Support the auto completion of while_op

* add dist_fill_constant_batch_size_like

* align infer  accuracy
```
  ceef73c9
- A
  [Eager] Add _fallback_legacy_dygraph for npu/xpu/rocm (#41774) · 5a103150
  由 Aurelius84 提交于 4月 18, 2022
```
* [Eager] add _fallback_legacy_dygraph for npu/xpu/rocm

* fix import
```
  5a103150
- Z
  
  Add sparse kernel coalesced (#41784) · 8f469ddd
  由 zhangkaihuo 提交于 4月 18, 2022
  
  8f469ddd
- S
  Optimization for graph_sample_neighbors API (#41447) · c31dd04c
  由 Siming Dai 提交于 4月 18, 2022
```
* add eids result for graph_sample_neighbors

* fix bug

* move fisher_yates sample to warp

* add cpu eid output

* delete comment

* delete comment

* change nullptr placeholder

* optimize sample kernel

* fix mutable_data
```
  c31dd04c
- Q
  [MLU]add op: reduce_sum, elementwise_sub (#41697) · 9f06069d
  由 qipengh 提交于 4月 18, 2022
```
* [MLU]add op: reduce_sum, elementwise_sub

* [MLU]del unrelated code
```
  9f06069d
17 4月, 2022 2 次提交

[Perf] Optimize dygraph scheduling performance (#41696) · 7ee31a96

由 Chen Weihang 提交于 4月 17, 2022

* split phi and fluid infermeta context

* resolve conflict

* fix type error

* optimize scheduling perf

* spec small vector size

* replace all grad var name

* fix test failed

* move init defalut signature

* polish details

* polish details

* fix no init bug

* init sig for tests

* add init sig for infer

* fix infrt error

* fix infrt failed

* fix kunlun error

* fix infrt failed

7ee31a96

[CustomOp] Fix PlaceType related compat error (#41826) · b5d9c31c

由 Chen Weihang 提交于 4月 17, 2022

* fix place type related compat error

* fix test failed

* remove dll decl

* revert place type change

* add dll decl

b5d9c31c

16 4月, 2022 3 次提交

B

fix_sharding_copy_right (#41849) · 5e5ae0a0
由 Baibaifan 提交于 4月 16, 2022

5e5ae0a0

Moe ref (#41864) · e9a63237

由 Roc 提交于 4月 16, 2022

* moe ref

* ref commit; test=document_fix

* update; test=document_fix

* update test=document_fix

* update; test=document_fix

e9a63237

Lml/prim op pywrapper (#41813) · ebf4fe6e

由 levi131 提交于 4月 16, 2022

* native commit for triple grad of sigmod

* Updated unittests files

* init functional jacobian api

* Updated trible_test func

* Updated gradient_checker & test_script

* finish test with dtype float32

* add float64 test case

* polish code

* use atol=1e-5 with dtype float64

* fix for ci

* set timeout for test_jacobian

* fix dygraph grad to support high differential

* polish API docstring

* Updated gradient checker and some related files

* fix double grad strip error for high differential

* fix double grad strip error for high differential

* Add Sigmoid triple grad tests

* fix dygraph double grad dtype error when calling for high differential senario

* Updated triple grad teses func

* Use np.random to initialize ddx

* Updated triple_grad_check func

* add todo for gradient checker and refine some comments

* remove additional code

* add test for warnging in backward.py

* format python code

* support multi input in triple gradient checker

* Add matmul triple grad kernel

* Updated comments of TODO

* Supported some special tests

* Change code-format to follow CI std

* Updated gradient_checker.py

* Fix conflicts

* Removed unnecessary printing log

* Change code style to follow CI std

* merge upstream

* add priops.py

* add_p

* rm useless files

* add sub_p mul_p div_p

* add sqrt_p and tanh_p

* add reshape_p

* add broadcast_p

* Add python primitive wrappers.

* Jvp rules updated.

* JVP rules done for all the 17 primops.

* quick check and fixes.

* add jvp(op, *args)

* add broadcast_p fill_constant_p matmul_p reduce_p reshape_p transpose_p

* add split_p and concat_p

* add gather_p and scatter_add_p

* add slice_select_p and slice_assign_p

* Add transpose rules.

* add multi input check for add_p, sub_p, mul_p, div_p

* update concat_p

* Linearize and transpose in progress..

* refine gather_p and scatter_add_p

* updated.

* update transpose.

* refine slice_assign_p and slice_select_p

* init commit for lower

* Merged with primitive ops.

* small update

* add rules for orig2prim and prim2orig

* add 9 test for prim ops

* add more test and fix some bug

* add more test

* register proto

* Adding primops test.

* add shape valid check for broadcast_p op, and add keepdim attr into reduce_p op proto

* support multi input and multi output for split_p and concat_p

* Test updated.

* update

* fix slice bug for slice_select_p and slice_assign_p

* updated.

* Ops updated.

* Refactor and bug fixes.

* updated.

* finish orig2prim and prim2orig rules

* dtype for axis attr should be long int

* update dtype for axis attr int64_t

* update for iscan CI

* Update primx.

* Refactor vars in primx.

* update for lower transform

* update primx.py

* update

* Fix linearize and transpose.

* Update is_dot

* Update is_dot

* Update is_dot

* add gradient aggregation, fix add_transpose.

* pass first linearize+transpose test.

* update test

* add_prim_op_pywrapper

* Add primops UT

* Fix set_value and update

* Fix code format and PR-CI-Coverage
Co-authored-by: Nveyron95 <veyron_wu@163.com>
Co-authored-by: NJiabin Yang <360788950@qq.com>
Co-authored-by: NTongxin Bai <waffle.bai@gmail.com>
Co-authored-by: N0x45f <wangzhen45@baidu.com>

ebf4fe6e

15 4月, 2022 16 次提交

Moe ref (#41836) · c37af19c

由 Roc 提交于 4月 15, 2022

* moe ref

* ref commit; test=document_fix

* update; test=document_fix

* update test=document_fix

c37af19c

[Yaml]add adamw yaml (#41678) · ea0a164b

由 chentianyu03 提交于 4月 15, 2022

* add adamw yaml

* fix test case error

* make the name of weight and bias in linear1 and linear2 to be constant

ea0a164b

[DoubleGrad] Enabled test_imperative_star_gan_with_gradient_penalty.py under eager mode (#41730) · 27f28e82

由 Zhanlue Yang 提交于 4月 15, 2022

* [DoubleGrad] Enabled double grad test cases in eager_mode for test_imperative_double_grad

* Fixed elementwise issue

* Addressed CI failures

* [DoubleGrad] Enabled test_imperative_triple_grad test cases under eager_mode

* [DoubleGrad] Enabled test_autograd_functional_dynamic.py under eager mode

* Enabled more test cases

* [DoubleGrad] Enabled test_imperative_star_gan_with_gradient_penalty.py under eager mode

* Adjusted test_imperative_star_gan_with_gradient_penalty.py

27f28e82

H
[Dygraph] Refactor Model Parallel in eager mode (#41761) · e6fb6599
由 Haohongxiang 提交于 4月 15, 2022
```
* refactor mp in eager mode

* update

* update

* add uts
```
e6fb6599
D
【GPUPS】add afsclient and gpupsutil (#41324) · 30a1213b
由 danleifeng 提交于 4月 15, 2022
```
* add gpupsutil and afsclient; test=develop
```
30a1213b
F

[MLU] add mlu softmax kernel (#41816) · 2d6b71a2
由 fwenguang 提交于 4月 15, 2022

2d6b71a2

Add eager string tensor (#41039) · a22b68b8

由 Jack Zhou 提交于 4月 15, 2022

* Add core.eager.StringTensor __init__ which pyarray args can be passed

* Add the numpy method of core.eager.StringTensor

* revert tensor.to_string modification

* Add ToPyObject for core.eager.StringTensor

* Add debug string for core.eager.StringTensor

* Remove place args of core.eager.StringTensor temporarily

* Fix check string_tensor error

* remove dtype of core.eager.StringTensor

* add core.eager.StringTensor unittest

* remove pstring from VarDesc

* Add InitStringTensorWithStringTensor

* Remove to_string modification

* Remove zero_copy arg from StringTensor creator

a22b68b8

A
[IPU] add mixed-precission support for ipu (#41733) · d7224482
由 Allen Guo 提交于 4月 15, 2022
```
* add mixed-precission support for ipu

* restore cast_model_to_fp16 api

* update UTs
```
d7224482
Z

Add API: Sparse Convolution3D (#41434) · 1665594d
由 zhangkaihuo 提交于 4月 15, 2022

1665594d

support no_need_buffer in eager_fluid state (#41720) · 840d2eb6

由 pangyoki 提交于 4月 15, 2022

* support no_need_buffer in eager_fluid state

* change no_need_buffer info from fwd_info to bwd_info

* fix CI fail, gru_unit donnot use no_need_buffer

* fix conflict between no_need_buffer and dispensable

* use tensor.define in dispensable

* solve conflict

* solve conflict

840d2eb6

A

【Hackathon No.25】为 Paddle 新增 nanquantile 数学计算API (#41343) · b9ee6a29
由 Asthestarsfalll 提交于 4月 15, 2022

b9ee6a29
Z

support KL2 multi-card training, refactor KL2 unittest, *test=kunlun (#41543) · 2eac4db8
由 zhangxiaoci 提交于 4月 15, 2022

2eac4db8

Change cuDNN Conv kernel for auto tune feature (#41313) · 35acfeda

由 limingshu 提交于 4月 15, 2022

* change cudnn helper for auto-tune

* Add FLAGS_use_autotune to set the global status of autotune and change the order of choosing algorithm.

* Fix the bug in calculating and printing current step cache hit rate.

* Improve the autotune cache and fix unittest.

* Change the key from AlgorithmType to int64_t.

* Fix unittest for cpu-only env.

* change ChooseAlgoByWorkspace for heuristic mode
Co-authored-by: NLiu Yiqun <liuyiqun01@baidu.com>

35acfeda

F

[MLU] add mlu activation kernels (#41751) · 10114859
由 fwenguang 提交于 4月 15, 2022

10114859
F
[MLU] add mlu new profiler (#41138) · fc208b7e
由 fwenguang 提交于 4月 15, 2022
```
* [MLU] add mlu new profiler

* fix format
```
fc208b7e
C
[Auto Parallel]update cluster (#41722) · 605552a9
由 caozhou 提交于 4月 15, 2022
```
* update cluster
```
605552a9

14 4月, 2022 6 次提交

C

fix dtype bug (#41802) · e7f0aa38
由 caozhou 提交于 4月 14, 2022

e7f0aa38
C

fix divide zero error when cpu only (#41794) · a4f3c0e9
由 chenjian 提交于 4月 14, 2022

a4f3c0e9
Z

Supplementary documents (#41700) · 64237c3f
由 zhangkaihuo 提交于 4月 14, 2022

64237c3f
L
executor perf statistics (#41648) · cbe7466f
由 liutiexing 提交于 4月 14, 2022
```
* executor perf statistics

* fix ut

* fix ut

* fix ut

* add ut

* add ut
```
cbe7466f

Fix to #38693 (minimal UT) (#41026) · d0f3296b

由 Jacek Czaja 提交于 4月 14, 2022

* Add UT

- Added missed data_layout

- Added missing conversions

- NDHWC added

- NDHWC support in data_transform

- another fix

- condddate change

- fix

u- fix

- fix

- fix

- fix

- fix

- fix to hack

- compilation fix

- fix to automatic merge

* - reduced UT

* - fix

* - lint

* - fix to lint

d0f3296b

FC+elementwise_add (residual connection) (#41776) · 92d8d0bc

由 Sławomir Siwek 提交于 4月 14, 2022

* Change tensor name to match activation

* declare fc_eltwise_add pass

* merge conv_eltwise refactor PR

* first compilable draft

* unittest feedback tools

* Fuse pass tester

* Move IsReachable() to shared file

* 100% coverage of fuse_pass_tester.cc

* register pass

* Add bias node

* Improve unit tests / remove bias node from pattern

* improve fc_eltwiseadd_unittest

* cancel eltwise_add fuse if act is already fused

* Add elementwise_input scale

* Residual MVP

* Add new FC attrs

* Add more test cases

* Add missing op attrs

* Adapt code to new Elementwise pattern

* reuse existing fcpattern

* improve code style

* remove unused arguments

* fix typo

* remove whitespace

* remove int8 related code

* Remove attributes from base ops

* style

* style check

* Remove input from base op

* Set attribute during fuse

* ut timeout

* download and test model

* DRY

* apply feedback from review

* Style check

* fix typo

* cosmetic changes

* explicitly set residual as output

* VIT-OCR accuracy check

* trigger CI

* remove whitespaces

* fix missing data file

92d8d0bc

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功