提交 · b18cbfb290054b915a66b266a7409aa0636bf9e1 · 机器未来 / Paddle

25 10月, 2021 2 次提交

add op: fused_feedforward(forward) (#35843) · b18cbfb2

由 zhangkaihuo 提交于 10月 25, 2021

这个PR只包含fused_feedforward前向的代码。

相关kernel实现：fused_dropout_act_bias, fused_residual_dropout_bias, fused_layernorm_residual_dropout_bias

fused_feedforward是一个融合算子，该算子对transformer模型的feed forward层的算子进行融合和封装，使得前端只呈现一个接口，通过融合减少部分访存和kernel launch的时间，以此提升性能。

b18cbfb2

J

fix pool2d convert case (#36667) · 087c3abe
由 JingZhuangzhuang 提交于 10月 24, 2021

087c3abe

24 10月, 2021 1 次提交
- Z
  
  Add the macro `-DPADDLE_WITH_CINN`. (#36660) · e2173b68
  由 Zhen Wang 提交于 10月 24, 2021
  
  e2173b68
22 10月, 2021 5 次提交

W
correct slice serialize data (#36588) · 5e880840
由 wenbin 提交于 10月 22, 2021
```
* slice

* add UT
```
5e880840
Z

add fp16 kernel for clip_op (#36577) · 1962d3af
由 zhangbo9674 提交于 10月 22, 2021

1962d3af

Fused attention op forward (#35905) · d4906214

由 Li Min 提交于 10月 22, 2021

功能：本PR的目标是提高attention模块的计算性能。
为了减少框架层对op的调度开销，本PR通过在C++层手动实现attention模块，对外提供attention 大op；
为了减少防存开销，本PR采取了两种优化方法：
（1）在q,k,v计算时通过共享输入X，将该处的gemm，transpose和bias add从三次调用减少为一次；
（2）使用kernel融合优化技术，在不同cuda kernel之间通过寄存器传输数据；

d4906214

[hapi] support dygraph amp O2 (#36441) · 08248db0

由 Leo Chen 提交于 10月 22, 2021

* [hapi] support dygrapg amp O2

* fix problem of static pure fp16 in hapi

* fix bug

* fix format

* fix ut

* follow comments

* update ut

* update amp save/load

* fix ut

* refine code format

08248db0

[PaddlePaddle Hackathon] add InceptionV3 (#36064) · ff06df6d

由 Nyakku Shigure 提交于 10月 22, 2021

* add inceptionv3
Co-authored-by: Ainavo <ainavo@163.com>
Co-authored-by: Npithygit <pyg20200403@163.com>

ff06df6d

21 10月, 2021 10 次提交

Z

[NPU] Add p_norm_grad (#36497) · ed478a3e
由 zhulei 提交于 10月 21, 2021

ed478a3e
R

add swish_op for npu (#36579) · 7eab0fa6
由 ronnywang 提交于 10月 21, 2021

7eab0fa6

Added matmul_v2+transpose+reshape fuse pass (#36481) · 856cb9c5

由 jakpiase 提交于 10月 21, 2021

* added base changes for matmul_v2+trans+resh fuse pass

* added full matmul_v2+transpose+reshape pass

* removed a file added by mistake

* added reviewers suggestions

* Changed ops type in checking capatibility version

* Deteled one statement

856cb9c5

[NPU] Add sync_batch_norm and sync_batch_norm_grad NPU Kernel (#36320) · 0ca2807c

由 furnace 提交于 10月 21, 2021

* add sync_batch_norm (support train, infer, and fp32, fp16, and NCHW, NHWC)

* [NPU] Delete debug codes

* [NPU] Remove FP16

0ca2807c

Add viterbi decode (#35778) · 6072aecb

由 Jack Zhou 提交于 10月 21, 2021

* add viterbi decode cpu kernel

* add viterbi decoder api in paddle.text

* add a data buffer once to avoid create many small pieces of data buffer frequently

* fix viterbi max_seq_length bug

* fix seq_len=1 bug

* fix device context

* move split out of for loop

* remove INVERSE_SUB

* remove 2 GET_CAST_MASK

* remove 1 loop

* remove Functor

* add to_static deploy code

* use MAX_FUNC instead of ELE_MAX

* add MaxFunctor

* impl max_func

* remove MaxFunctor

* remove cast op

* use REGISTER_OP_WITHOUT_GRADIENT

* add viterbi cuda kernel

* add FIX_BLOCKDIM_CASE macro

* add MKL add, mul; add get data mask

* add arange mkl impl

* add CPU Argmax

* add cpu gather

* use EXECUTE_MKL_ELEMENT_BINARY_OP instead of some ADD, MUL

* use SameDimsBinaryOP instead of EXECUTE_MKL_ELEMENT_BINARY_OP

* use SAME_DIMS_ELEMENT_BINARY_OP

* add SimpleBroadcastBinaryOP

* use int instead of int64_t to accelerate

* optimize SimpleBroadcastBinaryOP

* optimize SimpleBroadcastBinaryOP

* optimize performance in both single thread and multithread situation

* remove useless line

* remove useless code

* add CREATE_TENSOR_BUFFER macro

* add INIT_REQUIRED_TENSOR macro

* add comment

* fix windows ci

* add viterbi unittest

* remove cuda add functor

* remove cuda equal

* remove a template function

* fix windows ci

* fix windows dtype

* remove some template instance

* remove useless header file

* remove some blockdim

* remove transpose impl

* accelerate cpu performance on single thread situation

* viterbi_decode->crf_decode

* rename crf params name

* add viterbi api test

* remove useless import

* add enable_static

* use viterbi decoder

* fix viterbi len=1

* fix  viterbi unittest

* remove useless comments

* reconstruct viterbi decode

* remove ADD,SUB,MUL structure

* fix coverage

* remove CREATE_TENSOR

* add name args

* crf.py->ops.py; with_start_stop_tag->include_start_end_tag

* update crf_decode en docs

* fix viterbi decode en docs

* fix some review comments

* add FIXED_BLOCK_DIM_CASE in cuda

* push_back->emplace_back

* crf_decode->viterbi_decode; include_start_end_tag->include_bos_eos_tag

* paddle.text.ops.viterbi_decode->paddle.text.viterbi_decode

* fix viterbi_decode en docs

6072aecb

D

fix hdfs download_dir (#36590) · 66f4b292
由 danleifeng 提交于 10月 21, 2021

66f4b292
T
add fill_any_like/flatten ops to train ssd on kunlun (#36550) · 7bf2aa38
由 TTerror 提交于 10月 21, 2021
```
* add some ops to train ssd on kunlun

* update test_fill_any_like_op_xpu.py
```
7bf2aa38
X

User specified backend (#35745) · b6e7f8e9
由 xiongkun 提交于 10月 21, 2021

b6e7f8e9
Y

Fixed unit test for auto parallel cost model (#36574) · f6985774
由 YipZLF 提交于 10月 21, 2021

f6985774
Z

refine comments for GradScaler state_dict (#36522) · 1d38a013
由 zhangbo9674 提交于 10月 21, 2021

1d38a013

20 10月, 2021 10 次提交

H
fix bugs of ClipGradByGlobalNorm in HybridParallel (#36555) · 6a3941e3
由 Haohongxiang 提交于 10月 20, 2021
```
* fix bugs of ClipGradByGlobalNorm

* add unittests

* add unittests
```
6a3941e3
李
Fix global gather and global scatter operators (#36517) · 17b4dd70
由李季提交于 10月 20, 2021
```
* fix global gather and global scatter operators
```
17b4dd70
R

[NPU] Add kldiv_loss_op for npu (#36494) · 6a572a19
由 ronnywang 提交于 10月 20, 2021

6a572a19

Add FasterTokenizer Operator (#34491) · 3f2d6a3f

由 Steffy-zxf 提交于 10月 20, 2021

Add Tokenizer related functionalities for Transformer model in order that the process of training and predicting is consistent.

* support the text string as an input Tensor
* support the "VOCAB"unordered_map<wstring, int> as an input Tensor to lookup tokens
* Tokenizer used for BERT. This tokenizer applies an end-to-end, text string to wordpiece tokenization.
* It first applies basic tokenization, followed by wordpiece tokenization.

3f2d6a3f

Z

fix pow2 decay (#36559) · 605e7f08
由 Zeng Jinle 提交于 10月 20, 2021

605e7f08
W

add unittest (#36371) · 7325c9fb
由 Wilber 提交于 10月 20, 2021

7325c9fb
W

update for trt convert ut. (#36549) · 06bd348d
由 Wilber 提交于 10月 20, 2021

06bd348d

[FIX] Extend time for test_activation_nn_grad to avoid its timeout issue (#36527) · c285c719

由 Jiabin Yang 提交于 10月 20, 2021

* native commit for triple grad of sigmod

* Updated unittests files

* init functional jacobian api

* Updated trible_test func

* Updated gradient_checker & test_script

* finish test with dtype float32

* add float64 test case

* polish code

* use atol=1e-5 with dtype float64

* fix for ci

* set timeout for test_jacobian

* fix dygraph grad to support high differential

* polish API docstring

* Updated gradient checker and some related files

* fix double grad strip error for high differential

* fix double grad strip error for high differential

* Add Sigmoid triple grad tests

* fix dygraph double grad dtype error when calling for high differential senario

* Updated triple grad teses func

* Use np.random to initialize ddx

* Updated triple_grad_check func

* add todo for gradient checker and refine some comments

* remove additional code

* add test for warnging in backward.py

* add tanh triple grad

* format python code

* refine code

* make test_activation_nn_grad test time to 150s
Co-authored-by: Nveyron95 <veyron_wu@163.com>
Co-authored-by: Nlevi131 <limaolin01@baidu.com>

c285c719

[Auto Parallel] Generalization for Partition and Completion (#35735) · 797bd40d

由 JZ-LIANG 提交于 10月 20, 2021

* default dist op

* add dist_attr for dist op

* add unitest

* update inputname

* update function name

* add unitest

* update CMakeLists.txt for CI

* fix dis_matmul

* fix compile error

* update matmul to matmul_v2

* unify api

* unify api

* todo

* update distop forward func

* update distop forward func

* auto parallel backward

* update dist op

* autoparallel backward

* add backward for embedding

* temp1

* temp2

* temp3

* temp4

* backward done1

* backward done2

* backward done3

* dist embedding remove mp mode

* dist matmul remove mp mode

* update dist embedding
『

* dist op init1

* dist op init 2

* update unitest

* context remove parallel mode

* partitioner remove parallel mode

* update unitest

* a more general method to support varying mesh in pipeline parallel

* support varying mesh in pipeline parallel

* embedding support varying mesh in pipeline parallel

* matmul support varying mesh in pipeline parallel

* default dist op support varying mesh in pipeline parallel

* dist attribute for startup program

* default dist op support varying mesh in pipeline parallel 2

* partitoner support varying mesh in pipeline parallel

* revise logic for auto compeletion

* revise framework.py

* revise reshard unitest

* revise unitest for parallelize

* chmod

* fixed bug for dist embedding name mapping
Co-authored-by: Nzhaoyingli <zhaoyingli@baidu.com>

797bd40d

remove no_value using var.name (#36513) · fe01ba6a

由 0x45f 提交于 10月 20, 2021

* remove no_value using var.name

* fix unit test for CI

* fix unit test

* add test case

* fix test case

* add more test case

fe01ba6a

19 10月, 2021 12 次提交
- W
  Support elementwise_add triple grad Kernel (#36508) · 51c97d9f
  由 Weilong Wu 提交于 10月 19, 2021
```
* Support elementwise_add triple grad Kernel

* Change code-format to follow CI std
```
  51c97d9f
- Z
  [NPU] Add iou_similarity op (#36412) · 999242e3
  由 zhulei 提交于 10月 19, 2021
```
* [NPU] Add iou_similarity op

* [NPU] Add iou_similarity op

* [NPU] Add iou_similarity op
```
  999242e3
- K
  
  fix op_flops not define. test=develop (#36489) · f2612462
  由 Kaipeng Deng 提交于 10月 19, 2021
  
  f2612462
- D
  
  [heterps]edit shrink and unseenday logit for pslib (#36194) · 9e494472
  由 danleifeng 提交于 10月 19, 2021
  
  9e494472
- W
  Inference add type check in copy_from_cpu (#36429) · be6a8330
  由 Wilber 提交于 10月 19, 2021
```
* update

* fix ut error

* update ut
```
  be6a8330
- W
  add nearest_interp_v2 trt plugin (#34126) · 7b67f398
  由 wangxinxin08 提交于 10月 19, 2021
```
* add nearest_interp_v2 trt plugin
```
  7b67f398
- W
  
  [hybrid] static model parallel dropout support deterministic RandomSeedGenerator (#36228) · 8cc8e411
  由 WangXi 提交于 10月 19, 2021
  
  8cc8e411
- fix replicate pad when input size is 0 (#36510) · d89a759b
  由 littletomatodonkey 提交于 10月 19, 2021
```
* fix replicate pad when input size is 0
* add unit test
```
  d89a759b
- X
  catch the generatorfunction and intercept it. (#35369) · 7edcc4fb
  由 xiongkun 提交于 10月 19, 2021
```
* catch the generatorfunction and intercept it.

* add test generator

* add test case

* refine the testcase
```
  7edcc4fb
- Y
  [paddle.linalg.qr] Add the Qr Operator (#35742) · 34d785c2
  由 Yulong Ao 提交于 10月 19, 2021
```
* Add QR decomposition op

* Change codes to adapt to new svd_helper

* Update linalg.py

Restore the deleted comma

* Restore the deleted line

* Update linalg.py

* Update linalg.py

* Improve the qr code by reviews

* Update QR based on CI results

* Update qr doc, test=document_fix

* Change unsafe and ill-formed codes
```
  34d785c2
- Y
  Add auto parallel cost model and unittests (#36363) · a573a7ed
  由 YipZLF 提交于 10月 19, 2021
```
* Add auto parallel cost model and unittests

* Fixed code styles.

* Fixed bugs and codes style

* fixed typo

* Improved code style: object encapsulation.

* Fixed codes.

* Refractored estimate_cost

* Fixed typo
```
  a573a7ed
- X
  
  fix out of range for area interp, test=develop (#36466) · 77f4597f
  由 xiaoting 提交于 10月 19, 2021
  
  77f4597f

机器未来 / Paddle 与 Fork 源项目一致

机器未来 / Paddle
与 Fork 源项目一致