提交 · fb8be035ec9f4e55db4deb62cf9e6de8fabdb6eb · BaiXuePrincess / Paddle

22 9月, 2021 1 次提交
- [cherry-pick2.2]support extern third_party lapack API on Linux/Windows/Mac (#35897) · fb8be035
  由 zhouweiwei2014 提交于 9月 22, 2021
```
ATT, cherry-pick #35690
```
  fb8be035
18 9月, 2021 6 次提交

H

fix import paddle · edeb0ade
由 huangjun12 提交于 9月 17, 2021

edeb0ade
H

replace matmul to matmul_v2, expand to expand_v2 · 1776293a
由 huangjun12 提交于 9月 16, 2021

1776293a

由 Feiyu Chan 提交于 9月 18, 2021

* 1. add interface for fft;
2. add data type predicate;
3. fix paddle.roll.

* add fft c2c cufft kernel

* implement argument checking & op calling parts for fft_c2c and fftn_c2c

* add operator and opmaker definitions

* only register float and double for cpu.

* add common code for implementing FFT, add pocketfft as a dependency

* add fft c2c cufft kernel function

* fix bugs in python interface

* add support for c2r, r2c operators, op makers, kernels and kernel functors.

* test and fix bugs

* 1. fft_c2c function: add support for onesided=False;
2. add complex<float>, complex<double> support for concat and flip.

* 1. fft: fix python api bugs;
2. shape_op: add support for complex data types.

* fft c2c cufft kernel done with complie and link

* fix shape_op, add mkl placeholder

* remove mkl

* complete fft c2c in gpu

* 1. implement mkl-based fft, FFTC2CFunctor and common function exec_fft;
2. change the design, add input and output typename as template parameter for all FFTFunctors, update pocketfft-based implementation.

* complete fft c2c on gpu in ND

* complete fft c2c on gpu in ND

* complete fft c2c backward in ND

* fix MKL-based implementation

* Add frame op and CPU/GPU kernels.

* Add frame op forward unittest.

* Add frame op forward unittest.

* Remove axis parameter in FrameFunctor.

* Add frame op grad CPU/GPU kernels and unittest.

* Add frame op grad CPU/GPU kernels and unittest.

* Update doc string.

* Update after review and remove librosa requirement in unittest.

* Update grad kernel.

* add fft_c2r op

* Remove data allocation in TransCompute function.

* add fft r2c onesided with cpu(pocketfft/mkl) and gpu

* last fft c2r functor

* fix C2R and R2C for cufft, becase the direction is not an option in these cases.

* add fft r2c onesided with cpu(pocketfft/mkl) and gpu

* fix bugs in python APIs

* fix fft_c2r grad kernal

* fix bugs in python APIs

* add cuda fft c2r grad kernal functor

* clean code

* fix fft_c2r python API

* fill fft r2c result with conjugate symmetry (#19)

fill fft r2c result with conjugate symmetry

* add placeholder for unittests (#24)

* simple parameterize test function by auto generate test case from parm list (#25)

* miscellaneous fixes for python APIs (#26)

* add placeholder for unittests

* resize fft inputs before computation is n or s is provided.

* add complex kernels for pad and pad_grad

* simplify argument checking.

* add type promotion

* add int to float or complex promotion

* fix output data type for static mode

* fix fft's input dtype dispatch, import fft to paddle

* fix typos in axes checking (#27)

* fix typos in axes checking

* fix argument checking (#28)

* fix argument checking

* Add C2R Python layer normal and abnormal use cases (#29)

* documents and single case

* test c2r case

* New C2R Python layer normal and exception use cases

* complete rfft,rfft2,rfftn,ihfft,ihfft2,ihfftn unittest and doc string (#30)

* Documentation of the common interfaces of c2r and c2c (#31)

* Documentation of the common interfaces of c2r and c2c

* clean c++ code  (#32)

* clean code

* Add numpy-based implementation of spectral ops (#33)

* add numpy reference implementation of spectral ops

* Add fft_c2r numpy based implementation for unittest. (#34)

* add fft_c2r numpy implementation

* Add deframe op and stft/istft api. (#23)

* Add frame api

* Add deframe op and kernels.

* Add stft and istft apis.

* Add deframe api. Update stft and istft apis.

* Fix bug in frame_from_librosa function when input dims >= 3

* Rename deframe to overlap_add.

* Update istft.

* Update after code review.

* Add overlap_add op and stft/istft api unittest (#35)

* Add overlap_add op unittest.

* Register complex kernels of squeeze/unsquuze op.

* Add stft/istft api unittest.

* Add unittest for fft helper functions (#36)

* add unittests for fft helper functions. add complex kernel for roll op.

* complete static graph unittest for all public api (#37)

* Unittest of op with FFT C2C, C2R and r2c added (#38)

* documents and single case

* test c2r case

* New C2R Python layer normal and exception use cases

* Documentation of the common interfaces of c2r and c2c

* Unittest of op with FFT C2C, C2R and r2c added
Co-authored-by: lijiaqi <lijiaqi0612@163.com>

* add fft related options to CMakeLists.txt

* fix typos and clean code (#39)

* fix invisible character in mkl branch and fix error in error message

* clean code: remove docstring from unittest for signal.py.

* always convert numpy array to paddle.Tensor to avoid comparing numpy dtype with paddle dtype. (#40)

* always convert numpy array to paddle.Tensor to avoid comparing numpy dtype with paddle dtype.

* fix CI Errors: numpy dtype comparison, thrust when cuda is not available (#41)

1. always convert numpy array to paddle.Tensor to avoid comparing numpy dtype with paddle dtype.
2. promote floating point tensor to complex tensor ior fft_c2c and fft_c2r;
3. fix unittest to catch UnImplementedError and RuntimeError;
4. fix compile error by avoid using thrust when cuda is not available.
5.  fix sample code, use paddle.fft instead of paddle.tensor.fft

* remove inclusion of thrust, add __all__ list for fft (#42)

* Add api doc and update unittest. (#43)

* Add doc strings.
* Update overlap_add op unittest

* fix MKL-based FFT implementation (#44)

* fix MKL-based FFT implementation, MKL CDFT's FORWARD DOMAIN is always REAL for R2C and C2R

* remove code for debug (#45)

* use dynload for cufft (#46)

* use std::ptrdiff_t as datatype of stride (instead of int64_t) to avoid argument mismatch on some platforms.

* add complex support for fill_zeros_like

* use dynload for cufft

* Update doc and unittest. (#47)

* Add doc of frame op and overlap_add op.

* Update unittest.

* use dynload for cufft (#48)

1. use dynload for cufft
2. fix unittest;
3. temporarily disable Rocm.

* fix conflicts and merge upstream (#49)

fix conflicts and merge upstream

* fix compile error: only link dyload_cuda when cuda is available (#50)

* fix compile error: only link dyload_cuda when cuda is available

* fix dynload for cufft on windows (#51)

1. fix dynload for cufft on windows;
2. fix unittests.

* add NOMINMAX to compile on windows (#52)

 add NOMINMAX to compile on windows

* explicitly specify capture mode for lambdas (#55)

 explicitly specify capture mode for lambdas

* fix fft sample (#53)

* fix fft sample

* update scipy and numpy version for unittests of fft (#56)

update scipy and numpy version for unittests of fft

* Add static graph unittests of frame and overlap_add api. (#57)

* Remove cache of cuFFT & Disable ONEMKL (#59)

1. replace numpy.fft with scipy.fft as numpy<1.20 not support ortho norm
2. remove cache of cufft plans;
3. enhance error checking.
4. default WITH_ONEMKL to OFF
Co-authored-by: Njeff41404 <jeff41404@gmail.com>
Co-authored-by: Nroot <root@bjyz-sys-gpu-kongming9.bjyz.baidu.com>
Co-authored-by: NKP <109694228@qq.com>
Co-authored-by: lijiaqi <lijiaqi0612@163.com>
Co-authored-by: NXiaoxu Chen <chenxx_id@163.com>
Co-authored-by: Nlijiaqi0612 <33169170+lijiaqi0612@users.noreply.github.com>

11518a43

W

trt support serialize and deserialize (#35828) · ba71421c
由 Wilber 提交于 9月 18, 2021

ba71421c
A
Clean ParseMemInfo and Fix unittest failed under multi-thread (#35840) · 2fff5a58
由 Aurelius84 提交于 9月 18, 2021
```
* Clean ParaseMemInfo and fix unittest with multi-thread

* fix declare
```
2fff5a58

Add new API "eigvals" in linalg (#35720) · d411a038

由 From00 提交于 9月 18, 2021

* Add linalg.eigvals API

* pre-commit check

* Adjust code style

* Fix conflict

* Improve code style

* Modify the test code to ignore testing CUDA kernel

* Sort ouput data before checking in test code

* Set timeout value for UT

* Improve API example code to pass CI

* Fix bug for None fetch_list in Windows

* Delete grad Op

d411a038

17 9月, 2021 19 次提交

[AMP] Support pure fp16 training mode for dygraph (#35521) · adaeee4d

由 zhangbo9674 提交于 9月 17, 2021

* add pure fp16 major function in auto_cast & tracer

* support master weight in dygraph for pure fp16

* check mix dtype of fp16&fp32 for check_finite_and_unscale op

* change pure fp16 funtion name

* refine some bug in auto_cast

* refine auto_cast interface logic

* add param _casted_by_pure_fp16 for class Layer

* support state_dict hook for save model by user appointed dtype in pure_fp16_decorator

* refine pure_fp16_decorator as decorator

* add unittest

* add comment

* add comment

* support recompute

* add comment for auto_cast and decorator

* support to_static_state_dict for paddle.jit.save

* unlimite models num and optimizers num

* add lookup_table in black_list

* fix momentum and layer state_dict

* fix bug in layer state_dict

* fix bug in layer state_dict_helper

* refine unittest

* refine test_momentun_op

* refine interface and some code

* refine amp_decorator interface

* refine pure fp16 interface

* refine master weight interface

adaeee4d

L
temporally disable the warnings (#35560) · 68ae6345
由 Leo Chen 提交于 9月 17, 2021
```
* temporally disable the warnings

* disable ut
```
68ae6345
G

fix unittest (#35808) · fcfb0afe
由 Guoxia Wang 提交于 9月 17, 2021

fcfb0afe

Add linalg pinv api (#35804) · 71e01d3f

由 andyjpaddle 提交于 9月 17, 2021

* add pinv api, test=develop
* add linalg pinv api, test=develop
* update example code, test=develop

71e01d3f

Support EMA in Paddle2.x and Fleet (#35673) · fb4d5689

由 Haohongxiang 提交于 9月 17, 2021

* Support EMA in Paddle2.x and Fleet

* update

* update

* update

* modify ut of ema

* modify docs

* modify bugs

* update

* update

* update

* modify ut

fb4d5689

G

test=document_fix (#35824) · 177bf52f
由 Guoxia Wang 提交于 9月 17, 2021

177bf52f

add inplace op support to prune, scale_op is no longer need in jit.save (#35730) · 21921936

由 Haipeng Wang 提交于 9月 17, 2021

* add scale_op in model save step is not necessary, just fix the prune method to support static graph and inplace op

* fix jit.save, no need to add scale_op to each outputvar anymore.
fix prune_with_input, now it supports inplace op

* temporarily disable test_trt_dynamic_shape.TRTDynamicShapeOutOfBound2Test

21921936

津

[inference]add hard_swish dynamic plugin (#35214) · c59c8e4f
由津提交于 9月 17, 2021

c59c8e4f
Z

Fix segment api document. (#35818) · 6d5fc220
由 Zhong Hui 提交于 9月 17, 2021

6d5fc220

Add skip teller (#35807) · 0f74e5e7

由 xiaoxiaohehe001 提交于 9月 17, 2021

* add_skip_layernorm

* add_skip_layernorm

* add_skip_layernorm

* add_skip_layernorm

* add_skip_layernorm

* add_skip_layernorm

* add_skiplayernorm_teller

* add_skip_layernorm

* add_skip_layernorm_teller

* add_skip_layernorm_teller

* add_skip_layernorm

* add_skip_teller

0f74e5e7

L
expose cuda stream to users (#35813) · 40cfa512
由 Leo Chen 提交于 9月 17, 2021
```
* expose cuda stream to users

* add ut
```
40cfa512
津
[inference]add reduce converter test (#35145) · 05275010
由津提交于 9月 17, 2021
```
* add test

* add test

* add test
```
05275010
津
leaky_relu test (#35318) · 867f4fa0
由津提交于 9月 17, 2021
```
* add test

* add test

* add test

* add test

* add test
```
867f4fa0
W

polish code. (#35783) · 61010bb8
由 WeiXin 提交于 9月 17, 2021

61010bb8
X
fix unpool doc, test=document_fix (#35806) · 652e655f
由 xiaoting 提交于 9月 17, 2021
```
* fix unpool doc, test=document_fix

* fix typo for python example, test=document_fix
```
652e655f
J

update acc func using topk v2 (#35789) · 94d2cf82
由 Jiaqi Liu 提交于 9月 17, 2021

94d2cf82

增强equal API，输入Y支持int，float，bool或者tensor类型 (#35695) · 9b2d53fc

由 yeliang2258 提交于 9月 17, 2021

* update equal op, input Y can be float,int,bool or tensor

* update test

* update code style

* update code style

* update doc

* update str check

* remote str

* add type check

9b2d53fc

0

refine matrix_rank op code and doc (#35722) · 28fffef6
由 0x45f 提交于 9月 17, 2021

28fffef6
G
add launch doc (#35634) · 5548061b
由 Guoxia Wang 提交于 9月 17, 2021
```
* add launch doc
```
5548061b

16 9月, 2021 14 次提交

fix bug in pscore (#35698) · e64fed86

由 wangguanqun 提交于 9月 16, 2021

* add trainer desc config to distributed strategy

* code style modified

* data_feed set lod

* fix bug

* code style

* fix bug

e64fed86

Y

[hybrid] Fix mp multi gradient clip prob (#35713) · a4eadd15
由 Yuang Liu 提交于 9月 16, 2021

a4eadd15
Z

Add segment apis to paddle.incubate (#35759) · 4b683887
由 Zhong Hui 提交于 9月 16, 2021

4b683887
A
[NPU] add index_select_grad kernel and unit tests (#35594) · 67a094b5
由 Aganlengzi 提交于 9月 16, 2021
```
* [NPU] add index_select_grad kernel and unit tests

* dim=0 not need transpose
```
67a094b5
K
fix dataloader exit terminate error (#34501) · e93c18a3
由 Kaipeng Deng 提交于 9月 16, 2021
```
* fix DataLoader exit with SIGABRT/SIGSEGV. test=develop
```
e93c18a3

Support new API linalg.cond in paddle (#35140) · 2df74aa6

由 Haohongxiang 提交于 9月 16, 2021

* Support new API linalg.cond in paddle

* check code style

* check code style

* modify codes

* add docs_eng of linalg.cond

* add svd_norm for linalg.cond

* modify docs_en of cond

* add support for empty input in dynamic mode

* modify set_time of unittest

* update

* modify unittest of cond

* update

* remove cond in paddle.__all__

* pull latest codes

* merge latest codes

* update

2df74aa6

C

Add CPU and GPU eigh op implementation (#34990) · 07d0b834
由 crystal 提交于 9月 16, 2021

07d0b834

Remove autograd/grad api (#35579) · 7d9ca164

由 chentianyu03 提交于 9月 16, 2021

* remove autograd/grad api

* import grad, no_grad set_grad_enable from autograd module

* modify import no_grad_ as no_grad

7d9ca164

L

add pretrained weight of vgg19 (#35788) · 7ec90060
由 LielinJiang 提交于 9月 16, 2021

7ec90060
W
[paddle-trt] fix gather convert (#35784) · 7546a079
由 Wangzheee 提交于 9月 16, 2021
```
* fix gather

* fix
```
7546a079

[Dy2stat]fix no_grad context error in dy2stat (#35725) · 3e897489

由 0x45f 提交于 9月 16, 2021

* fix no_grad context error in dy2stat

* remove useless comments

* fix error by drop_kids in python

* add test and fix review

3e897489

G
support l2_normalize float16 (#35776) · b666fd3c
由 Guoxia Wang 提交于 9月 16, 2021
```
* support fp16 dtype
```
b666fd3c
L
remove distributed attributes at the last stage for auto parallel (#35605) · a3790606
由 lilong12 提交于 9月 16, 2021
```
* update
```
a3790606

Python support register pass via PassDesc (#35602) · bab39eb2

由 wuhuanzhou 提交于 9月 16, 2021

PR主要功能：针对fusion等子图替换场景，支持Python侧开发并注册Pass。

背景
Pass是指输入一个深度学习计算图Graph，依照一定条件进行修改，输出修改后的Graph的过程；
当前PaddlePadle框架编写Pass代码存在以下问题：
用户需要手写Graph的条件匹配、在Graph上的修改代码；
对Graph操作需要深入底层框架代码，了解Graph的结构，并且知道相关Pass写法；
我们提出了针对fusion等子图替换类Pass的优化方案以支持用户在Python侧开发注册Pass，提升二次开发体验：
用户只需要输入匹配和替换的子图描述，由深度学习框架编写的代码来生成匹配和替换的逻辑，不需要用户对Graph进行匹配和替换操作；
API级别的替换，用户可以通过Paddle的Python API构造子图，从而不需要知道Graph的结构，也能写Paddle的Graph Pass代码

bab39eb2

BaiXuePrincess / Paddle 与 Fork 源项目一致

BaiXuePrincess / Paddle
与 Fork 源项目一致