提交 · d4906214656754cb7216279d07b327b49c3ee617 · 机器未来 / Paddle

22 10月, 2021 5 次提交

Fused attention op forward (#35905) · d4906214

由 Li Min 提交于 10月 22, 2021

功能：本PR的目标是提高attention模块的计算性能。
为了减少框架层对op的调度开销，本PR通过在C++层手动实现attention模块，对外提供attention 大op；
为了减少防存开销，本PR采取了两种优化方法：
（1）在q,k,v计算时通过共享输入X，将该处的gemm，transpose和bias add从三次调用减少为一次；
（2）使用kernel融合优化技术，在不同cuda kernel之间通过寄存器传输数据；

d4906214

[hapi] support dygraph amp O2 (#36441) · 08248db0

由 Leo Chen 提交于 10月 22, 2021

* [hapi] support dygrapg amp O2

* fix problem of static pure fp16 in hapi

* fix bug

* fix format

* fix ut

* follow comments

* update ut

* update amp save/load

* fix ut

* refine code format

08248db0

【Bug Fixes】Elementwise_add triple grad, fixed an input uninitialized problem (#36618) · 6580ad16

由 Weilong Wu 提交于 10月 22, 2021

* Support elementwise_add triple grad Kernel

* Change code-format to follow CI std

* Removed unreasonable code, and fixed an input uninitialized issue

* Support elementwise_add triple grad Kernel

* Change code-format to follow CI std

* Removed unreasonable code, and fixed an input uninitialized issue

6580ad16

W

support lite xpu choose device id (#36610) · f46311b0
由 Wilber 提交于 10月 22, 2021

f46311b0

[PaddlePaddle Hackathon] add InceptionV3 (#36064) · ff06df6d

由 Nyakku Shigure 提交于 10月 22, 2021

* add inceptionv3
Co-authored-by: Ainavo <ainavo@163.com>
Co-authored-by: Npithygit <pyg20200403@163.com>

ff06df6d

21 10月, 2021 15 次提交

Z

[NPU] Add p_norm_grad (#36497) · ed478a3e
由 zhulei 提交于 10月 21, 2021

ed478a3e
R

add swish_op for npu (#36579) · 7eab0fa6
由 ronnywang 提交于 10月 21, 2021

7eab0fa6

Added matmul_v2+transpose+reshape fuse pass (#36481) · 856cb9c5

由 jakpiase 提交于 10月 21, 2021

* added base changes for matmul_v2+trans+resh fuse pass

* added full matmul_v2+transpose+reshape pass

* removed a file added by mistake

* added reviewers suggestions

* Changed ops type in checking capatibility version

* Deteled one statement

856cb9c5

[NPU] Add sync_batch_norm and sync_batch_norm_grad NPU Kernel (#36320) · 0ca2807c

由 furnace 提交于 10月 21, 2021

* add sync_batch_norm (support train, infer, and fp32, fp16, and NCHW, NHWC)

* [NPU] Delete debug codes

* [NPU] Remove FP16

0ca2807c

Add viterbi decode (#35778) · 6072aecb

由 Jack Zhou 提交于 10月 21, 2021

* add viterbi decode cpu kernel

* add viterbi decoder api in paddle.text

* add a data buffer once to avoid create many small pieces of data buffer frequently

* fix viterbi max_seq_length bug

* fix seq_len=1 bug

* fix device context

* move split out of for loop

* remove INVERSE_SUB

* remove 2 GET_CAST_MASK

* remove 1 loop

* remove Functor

* add to_static deploy code

* use MAX_FUNC instead of ELE_MAX

* add MaxFunctor

* impl max_func

* remove MaxFunctor

* remove cast op

* use REGISTER_OP_WITHOUT_GRADIENT

* add viterbi cuda kernel

* add FIX_BLOCKDIM_CASE macro

* add MKL add, mul; add get data mask

* add arange mkl impl

* add CPU Argmax

* add cpu gather

* use EXECUTE_MKL_ELEMENT_BINARY_OP instead of some ADD, MUL

* use SameDimsBinaryOP instead of EXECUTE_MKL_ELEMENT_BINARY_OP

* use SAME_DIMS_ELEMENT_BINARY_OP

* add SimpleBroadcastBinaryOP

* use int instead of int64_t to accelerate

* optimize SimpleBroadcastBinaryOP

* optimize SimpleBroadcastBinaryOP

* optimize performance in both single thread and multithread situation

* remove useless line

* remove useless code

* add CREATE_TENSOR_BUFFER macro

* add INIT_REQUIRED_TENSOR macro

* add comment

* fix windows ci

* add viterbi unittest

* remove cuda add functor

* remove cuda equal

* remove a template function

* fix windows ci

* fix windows dtype

* remove some template instance

* remove useless header file

* remove some blockdim

* remove transpose impl

* accelerate cpu performance on single thread situation

* viterbi_decode->crf_decode

* rename crf params name

* add viterbi api test

* remove useless import

* add enable_static

* use viterbi decoder

* fix viterbi len=1

* fix  viterbi unittest

* remove useless comments

* reconstruct viterbi decode

* remove ADD,SUB,MUL structure

* fix coverage

* remove CREATE_TENSOR

* add name args

* crf.py->ops.py; with_start_stop_tag->include_start_end_tag

* update crf_decode en docs

* fix viterbi decode en docs

* fix some review comments

* add FIXED_BLOCK_DIM_CASE in cuda

* push_back->emplace_back

* crf_decode->viterbi_decode; include_start_end_tag->include_bos_eos_tag

* paddle.text.ops.viterbi_decode->paddle.text.viterbi_decode

* fix viterbi_decode en docs

6072aecb

D

fix hdfs download_dir (#36590) · 66f4b292
由 danleifeng 提交于 10月 21, 2021

66f4b292
T
add fill_any_like/flatten ops to train ssd on kunlun (#36550) · 7bf2aa38
由 TTerror 提交于 10月 21, 2021
```
* add some ops to train ssd on kunlun

* update test_fill_any_like_op_xpu.py
```
7bf2aa38
X

User specified backend (#35745) · b6e7f8e9
由 xiongkun 提交于 10月 21, 2021

b6e7f8e9

Fix a bug in ReadData, ReadDataBc and ReadDataReduce when NX != 1 (#36373) · 921c0917

由 niuliling123 提交于 10月 21, 2021

* Update the implement of reduceAnyKernel according to kernel primitive api
* Fix a bug in ReadData, ReadDataBc and ReadDataReduce when NX != 1

921c0917

S

Graph engine4 (#36587) · 5eb640c6
由 seemingwang 提交于 10月 21, 2021

5eb640c6
Z
add ctr table depends (#36465) · d64f7b3b
由 zhaocaibei123 提交于 10月 21, 2021
```
* add ctr table depends

* code style

* fix

* fix

* fix naming

* rename

* rename
```
d64f7b3b

Fix flame graph (#36578) · 72533986

由 liutiexing 提交于 10月 21, 2021

* add align for WorkQueue

* add spinlock

* merge develop

* merge

* Add EventsWaiter

* Revert "Add EventsWaiter"

This reverts commit e206173aa9be7401b83a53581627bfaf557c8fb2.

* adjust multithread using, fix flame graph

* update

72533986

Y

Fixed unit test for auto parallel cost model (#36574) · f6985774
由 YipZLF 提交于 10月 21, 2021

f6985774
Z

refine comments for GradScaler state_dict (#36522) · 1d38a013
由 zhangbo9674 提交于 10月 21, 2021

1d38a013
A
Support No DataTransform From GetKernelTypeForVar (#36571) · e82c3a5f
由 Aurelius84 提交于 10月 21, 2021
```
* Add kQueueSync.synchronize_run_ logic

* Support No DataTransform From GetKernelTypeForVar
```
e82c3a5f

20 10月, 2021 17 次提交

[heterps]fix heterps pipeline training (#36512) · ded3e705

由 danleifeng 提交于 10月 20, 2021

* split into PreBuildTask and BuildPull; slove endpass bug;test=develop

* change buildcpu into prebuild and buildcpu into build;test=develop

ded3e705

H
fix bugs of ClipGradByGlobalNorm in HybridParallel (#36555) · 6a3941e3
由 Haohongxiang 提交于 10月 20, 2021
```
* fix bugs of ClipGradByGlobalNorm

* add unittests

* add unittests
```
6a3941e3
李
Fix global gather and global scatter operators (#36517) · 17b4dd70
由李季提交于 10月 20, 2021
```
* fix global gather and global scatter operators
```
17b4dd70
R

[NPU] Add kldiv_loss_op for npu (#36494) · 6a572a19
由 ronnywang 提交于 10月 20, 2021

6a572a19
W

fix fc fuse proble (#36568) · fc5db55a
由 Wilber 提交于 10月 20, 2021

fc5db55a

Add FasterTokenizer Operator (#34491) · 3f2d6a3f

由 Steffy-zxf 提交于 10月 20, 2021

Add Tokenizer related functionalities for Transformer model in order that the process of training and predicting is consistent.

* support the text string as an input Tensor
* support the "VOCAB"unordered_map<wstring, int> as an input Tensor to lookup tokens
* Tokenizer used for BERT. This tokenizer applies an end-to-end, text string to wordpiece tokenization.
* It first applies basic tokenization, followed by wordpiece tokenization.

3f2d6a3f

W

adapt to cann5.0.3_alpha3. (#36106) · 873ee4e3
由 wuhuachaocoding 提交于 10月 20, 2021

873ee4e3
Z

fix pow2 decay (#36559) · 605e7f08
由 Zeng Jinle 提交于 10月 20, 2021

605e7f08
W

add unittest (#36371) · 7325c9fb
由 Wilber 提交于 10月 20, 2021

7325c9fb
W

update for trt convert ut. (#36549) · 06bd348d
由 Wilber 提交于 10月 20, 2021

06bd348d

fix SerializeSelectedRows (#36543) · 8ca5206b

由 zmx 提交于 10月 20, 2021

* bug fix for  DeserializeSelectedRows. test=develop

* fix bug for SerializeSelectedRows. test=develop

* update. test=develop

8ca5206b

Add CINN Compile Option (#36292) · 6524fa8d

由 Huihuang Zheng 提交于 10月 20, 2021

Add CINN compile option in CMake.

Now you can use CINN in Paddle by `-DWITH_CINN=ON` when `cmake`

To test it, you can run `make cinn_lib_test -j` and `ctest -R cinn_lib_test`. 

Note:
1. You should set
```
export runtime_include_dir=${CINN_SOURCE_DIR}/cinn/runtime/cuda 
```
When run test, the `${CINN_SOURCE_DIR}` should be set based on your CINN directory.

2. CINN is under developing now, you may have to change `CINN_GIT_TAG` to the git commit you need.

6524fa8d

W
fix (#36557) · 4bd19770
由 wenbin 提交于 10月 20, 2021
```
* fix

* remove const
```
4bd19770

[FIX] Extend time for test_activation_nn_grad to avoid its timeout issue (#36527) · c285c719

由 Jiabin Yang 提交于 10月 20, 2021

* native commit for triple grad of sigmod

* Updated unittests files

* init functional jacobian api

* Updated trible_test func

* Updated gradient_checker & test_script

* finish test with dtype float32

* add float64 test case

* polish code

* use atol=1e-5 with dtype float64

* fix for ci

* set timeout for test_jacobian

* fix dygraph grad to support high differential

* polish API docstring

* Updated gradient checker and some related files

* fix double grad strip error for high differential

* fix double grad strip error for high differential

* Add Sigmoid triple grad tests

* fix dygraph double grad dtype error when calling for high differential senario

* Updated triple grad teses func

* Use np.random to initialize ddx

* Updated triple_grad_check func

* add todo for gradient checker and refine some comments

* remove additional code

* add test for warnging in backward.py

* add tanh triple grad

* format python code

* refine code

* make test_activation_nn_grad test time to 150s
Co-authored-by: Nveyron95 <veyron_wu@163.com>
Co-authored-by: Nlevi131 <limaolin01@baidu.com>

c285c719

[Auto Parallel] Generalization for Partition and Completion (#35735) · 797bd40d

由 JZ-LIANG 提交于 10月 20, 2021

* default dist op

* add dist_attr for dist op

* add unitest

* update inputname

* update function name

* add unitest

* update CMakeLists.txt for CI

* fix dis_matmul

* fix compile error

* update matmul to matmul_v2

* unify api

* unify api

* todo

* update distop forward func

* update distop forward func

* auto parallel backward

* update dist op

* autoparallel backward

* add backward for embedding

* temp1

* temp2

* temp3

* temp4

* backward done1

* backward done2

* backward done3

* dist embedding remove mp mode

* dist matmul remove mp mode

* update dist embedding
『

* dist op init1

* dist op init 2

* update unitest

* context remove parallel mode

* partitioner remove parallel mode

* update unitest

* a more general method to support varying mesh in pipeline parallel

* support varying mesh in pipeline parallel

* embedding support varying mesh in pipeline parallel

* matmul support varying mesh in pipeline parallel

* default dist op support varying mesh in pipeline parallel

* dist attribute for startup program

* default dist op support varying mesh in pipeline parallel 2

* partitoner support varying mesh in pipeline parallel

* revise logic for auto compeletion

* revise framework.py

* revise reshard unitest

* revise unitest for parallelize

* chmod

* fixed bug for dist embedding name mapping
Co-authored-by: Nzhaoyingli <zhaoyingli@baidu.com>

797bd40d

A

Add kQueueSync.synchronize_run_ logic (#36546) · 127488ba
由 Aurelius84 提交于 10月 20, 2021

127488ba

remove no_value using var.name (#36513) · fe01ba6a

由 0x45f 提交于 10月 20, 2021

* remove no_value using var.name

* fix unit test for CI

* fix unit test

* add test case

* fix test case

* add more test case

fe01ba6a

19 10月, 2021 3 次提交
- W
  Support elementwise_add triple grad Kernel (#36508) · 51c97d9f
  由 Weilong Wu 提交于 10月 19, 2021
```
* Support elementwise_add triple grad Kernel

* Change code-format to follow CI std
```
  51c97d9f
- Z
  [NPU] Add iou_similarity op (#36412) · 999242e3
  由 zhulei 提交于 10月 19, 2021
```
* [NPU] Add iou_similarity op

* [NPU] Add iou_similarity op

* [NPU] Add iou_similarity op
```
  999242e3
- K
  
  fix op_flops not define. test=develop (#36489) · f2612462
  由 Kaipeng Deng 提交于 10月 19, 2021
  
  f2612462

机器未来 / Paddle 与 Fork 源项目一致

机器未来 / Paddle
与 Fork 源项目一致