提交 · 0c037d2d5918a4533764cc17dcd1cafd25aefa81 · wmsofts / Paddle

15 4月, 2021 2 次提交

F
fix test sync_with_cpp (#32212) · 0c037d2d
由 fangshuixun007 提交于 4月 15, 2021
```
fix test sync_with_cpp (#32212)
```
0c037d2d

【NPU】Cherry-pick ascendrc ops code by 0325 to develop (#32197) · e6bc358d

由 zhang wenhui 提交于 4月 15, 2021

* merge 31065

* Fix typo of selected_npus (#31230)

* merge 31249

* [NPU] Support npu op pow and pow grad (#31247)

* [NPU] Support npu op: (1) pow (2) pow_grad

* Support fp16

* Fix pow npu fp16 test (#31256)

* support list of list attribute for NPU (#31299)

* support list of list attribute for NPU

* fix compile problem

* fix reference

* [NPU] Support npu op: (1) slice (2) slice_grad (#31275)

* fix reading flags from env (#31329)

* merge 31347

* [NPU] Support npu op layer_norm and layer_norm_grad (#31310)

* init commit, add layer_norm npu kernel

* fix typo

* add unittest

* add unittest

* fix bug

* fix bug

* refine ut

* [NPU] add npu kernel for equal op (#31393)

* add npu kernel for equal op

* refine code

* add more ut

* update year

* [NPU] Support npu kernel for shape op  (#31427)

* add shape npu

* fix

* fix

* fix endif (#31431)

* Fix pow, use fillD instead of broadcast (#31433)

* Fix pow, refine code (#31440)

* fix cmake of cryptopp to avoid downloading every time (#31451)

* [NPU] squeeze and unsqueeze op for ascend (#31452)
Co-authored-by: Nroot <xiayanming@baidu.com>

* Support npu kernel for gather op (#31458)

* add gather npu op

* code review done

* update python new line

* precommit

* fix review

* del commit

* 【NPU】add scale op for npu (#31499)

* add scale npu

* fix

* fix

* Support TensorFormVector, TensorToVector of bool type (#31518)

* support TensorFormVector, TensorToVector of bool type

* add ut

* fix compile problem

* 【NPU】support npu kernel for fill_constant op (#31521)

* add fill_constant npu

* add fill_constant npu

* fix

* cherry-pick 31422, solve conflict

* 【NPU】Support npu kernel for matmul op (#31544)

* add matmulv2_npu

* add matmul

* add matmul

* [NPU] Support npu op elementwise_mul and elementwise_mul_grad (#31571)

* [NPU] Support npu op elementwise_max (#31574)

* 【NPU】add relu op for  npu (#31515)

* add relu npu

* fixed

* fix

* 【NPU】Suppert npu kernel for reshape2 op (#31524)

* add reshape2 npu

* add reshpe2

* [NPU] Support npu kernel for gather op fix bug (#31541)

* add gather npu op

* code review done

* update python new line

* precommit

* fix review

* del commit

* update gather_grad

* fix bug

* fix bug

* [NPU] Support npu kernel for amp_check_finite_and_unscale_npu op (#31457)

* Support npu kernel for amp_check_finite_and_unscale_npu op

* support EnforceNotMet exception

* fix exception bug

* modify python unittest

* precommit

* update c++ unittest

* fix review

* fix review

* [NPU] accuracy op (#31492)

* accuracy op

* fix license

* fix

* add test and fix bug

* [NPU] add Assign OP (#31561)

* add assign op

* add test assign npu test

* dele if def
Co-authored-by: Noyjxer <1728722986@qq.com>

* [NPU] fix npu op elementwise_mul_grad (#31592)

* 【NPU】Support npu op gelu and gelu_grad (#31530)

* Support npu op gelu and gelu_grad

* Support npu op gelu and gelu_grad

* [NPU] fix assgin cmake (#31595)

* fix gather_grad bug (#31607)

* [NPU] add range op (#31560)

* add range op

* fix codestyle; call GetSize directly
Co-authored-by: Noyjxer <1728722986@qq.com>

* 【NPU】Support npu op elementwise_div and elementwise_div_grad (#31573)

* Support npu op elementwise_div and elementwise_div_grad

* Support npu op elementwise_div and elementwise_div_grad

* Support npu op elementwise_div and elementwise_div_grad

* [NPU] Support npu op log, log_grad, sqrt, sqrt_grad, square, tanh and tanh_grad (#31600)

* [NPU] Support npu op logicalnot_op (#31534)

* [NPU] Support npu op elementwise_min (#31575)

* [NPU] Support npu op elementwise_pow (#31576)

* [NPU] Support npu op table_lookup_v2 and table_lookup_v2_grad (#31399)

* [npu] support npu kernel `table_lookup_v2`

* clean up

* +python test

* +cmake

* clean up

* remove int8 kernel
+ python unitest for fp16

* clean up

* [NPU] support npu kernel for `less_than` (#31327)

* [npu] support npu kernel for `less than`

* remove int* kernel

* cleanup

* [NPU] Support npu kernel scatter op (#31624)

* Support npu kernel scatter op

* Add more test

* [NPU] fix allocator min chunk size (#31632)

* [NPU] Support NPU kernel cast op (#31635)
Co-authored-by: Nfrankwhzhang <frankwhzhang@126.com>

* [NPU] add npu kernel for sgd (#31639)

* 【NPU】Support NPU kernel for reduce_sum op v2 (#31620)

* add reduce_sum

* fix broadcastd

* fix test

* fix

* add unsqueeze in reduce_sum

* add template

* add unittest for keep_dim

* test reduce_all
Co-authored-by: Nfrankwhzhang <frankwhzhang@126.com>

* [NPU] add npu kernel for adam (#31644)

* add npu kernel for adam

* refine code

* disable test

* modify atol

* 【NPU】Support npu kernel for mul op (#31584)

* add mul

* add test mul

* [NPU] add npu kernel for softmax_with_cross_entropy (#31656)

* init

* fix bugs

* [NPU] add npu kernel for mean Op (#31562)

* update mean op

* update mean op

* give a better test activation
Co-authored-by: Noyjxer <1728722986@qq.com>

* Revert "[NPU] add npu kernel for mean Op (#31562)" (#31665)

This reverts commit 468ac699.

* 【NPU】Add TensorCopy to NPU kernel for reduce_sum op  (#31667)

* update unittest

* add TensorCopy in npu grad kernel

* [NPU] Support npu op `expand` (#31405)

* [npu] support npu kernel  for `expand`

* [NPU] fix shape of dx in mul_grad (#31675)

* fix shape of dx

* refine code

* [NPU] add Increment op (#31563)

* add increment

* fix

* update test increment op inplace

* update increment op

* increment b = 2
Co-authored-by: Noyjxer <1728722986@qq.com>

* [NPU] add NPU add topk  (#31596)

* add topk op

* add cmake

* update topk npu op

* refactor func

* fix test not go npu TopKD bug

* NPUPlace(4) to NPUPlace(0)

* update comment
Co-authored-by: Noyjxer <1728722986@qq.com>

* [NPU] Support NPU kernel sum op (#31671)

* [NPU] npu support `transpose` (#31486)

* cherry-pick 31564, solve conflict

* [NPU] Fix bug: Fix calculation errors of pow grad npu kernel (#31699)

* [NPU] Support testing grad of NPU ops in OpTest (#31697)

* [NPU] Support NPU kernel of stack op (#31711)

* [NPU] Remove redundant ctest of top_k_op_npu_test (#31718)

* [NPU] fix reshape npu op kernel (#31726)

* rename npu op file

* fix reshape

* [NPU] change transpose to transpose2 (#31734)

* change transpose to transpose2

* fix bug

* [NPU] Support  mean npu kernel (#31729)

* [NPU] fix some bugs of npu op (#31739)

* fix softmax

* fix mean

* fix lookup_table_v2

* 【NPU】Fix npu kernel elementwise_div_grad  (#31753)

* [NPU] fix the grad kernel diff bug of gather op (#31757)

* fix gather grad kernel diff

* fix gather grad kernel diff

* fix gather review bug

* 【NPU】Fix reshape test & add grad test (#31776)

* fix

* fix

* [NPU] support fp16 for npu accuracy op (#31797)

* [NPU] support list of tensor input (#31801)

* support list of tensor as npu input

* add comment

* fix typo

* fix typo

* [NPU] add npu kernel for concat op (#31695)

* add npu kernel for concat op

* add npu kernel for concat op

* refine code

* update

* refine concat_grad

* [NPU] Support npu kernel for op elementwise_floordiv (#31822)

* [NPU] fix bug of lookup_table_v2_grad (#31834)

* [NPU] support default stream (#31510)

* [NPU] support mixed precision input for npu layer norm (#31847)

* support mixed precision input for npu layer norm

* fix layer_norm npu kernel
Co-authored-by: Nzhiqiu <chenqiuliang@baidu.com>

* 【NPU】Support npu kernel for update_loss_scaling op (#31830)

* add update_loss_scaling_npu NPU kernel

* change TensorFromVec to Memset

* fix compile problem (#31850)

* [NPU] support npu for conditional_block op (#31854)

* 【NPU】Add int dtype kernel for reshape2 op (#31864)

* fix

* fix

* [NPU] fix some op bugs (#31855)

* fix some op bugs

* fix some bugs

* follow comments

* fix log level

* add ut

* [NPU] support fp16 of input for api pow (#31871)

* [NPU] add npu kernel for truncated_gaussian_random op (#31654)

* init

* add todo

* add npu kernel for truncated_gaussian_random

* add sync

* fix concat_grad

* fix typo

* fix compile

* fix compile

* fix compile

* fix compile

* fix compile

* fix compile

* fix code style

* fix code style

* fix code

* Fix op test (#32231)

* fix conditional block (#32243)

* fix style code
Co-authored-by: Nxiayanming <41795079@qq.com>
Co-authored-by: NLeo Chen <chenqiuliang@baidu.com>
Co-authored-by: Nliym27 <33742067+liym27@users.noreply.github.com>
Co-authored-by: NReventon_L <luyuxiang1994@qq.com>
Co-authored-by: Nroot <xiayanming@baidu.com>
Co-authored-by: Noyjxer <1728722986@qq.com>
Co-authored-by: Nyinhaofeng <66763551+yinhaofeng@users.noreply.github.com>
Co-authored-by: NOleNet <olenet@126.com>
Co-authored-by: NMeiyim <chen_xuyi@outlook.com>
Co-authored-by: Noyxuan-11 <963650125@qq.com>
Co-authored-by: Npangyoki <pangyoki@126.com>

e6bc358d

14 4月, 2021 16 次提交
- Y
  
  Optimize the bec_loss op to avoid copy input back to CPU. (#32265) · 69d80274
  由 Yiqun Liu 提交于 4月 14, 2021
  
  69d80274
- A
  
  Optimize of backward of log_softmax when axis is -1 and dim_size <= 1024 (#32180) · 5dc0a6eb
  由 AshburnLee 提交于 4月 14, 2021
  
  5dc0a6eb
- W
  
  support the bool tensor and scalar (#32272) · 7da4455f
  由 wawltor 提交于 4月 14, 2021
  
  7da4455f
- J
  
  Added oneDNN reduce_op FWD kernel (#31816) · 3a804a0e
  由 jakpiase 提交于 4月 14, 2021
  
  3a804a0e
- A
  adds new CPU kernel for SGD op supporting BF16 data type (#32162) · 3ac6c189
  由 Adam Osewski 提交于 4月 14, 2021
```
* Initial draft for SGD BG16 kernel.

* Unit tests for SGD with BF16 data type.

* Add VLOG message to SGD BF16 op CPU kernel.

* Enhance error messages and error types.

* Refactor SGD op kernels to leverage some common code.

* Make easier to add new kerne invoke code.

* Fix SGD op kernel for sparse grad.

* Unify quotes style.

* Fix error for ROCM compilation.

* Use specialized PADDLE_ENFORCE_xx functions.
```
  3ac6c189
- C
  
  add marco cond for multi function (#32239) · 7b9fcaca
  由 Chen Weihang 提交于 4月 14, 2021
  
  7b9fcaca
- X
  
  softmax reconstruction and optimization (#31821) · 63abd500
  由 xingfeng01 提交于 4月 14, 2021
  
  63abd500
- P
  [Paddle-TRT] Add check for TRT runtime dynamic shape (#32155) · 8552a182
  由 Pei Yang 提交于 4月 14, 2021
```
* add check for runtime dynamic shape

* add unittest

* add lower bound case

* adjust timeout of new ut to 120s
```
  8552a182
- C
  Add inner register backward hook method for Tensor (#32171) · 7ba85aca
  由 Chen Weihang 提交于 4月 14, 2021
```
* add register backward hook method

* add leaf grad accumullated test
```
  7ba85aca
- Q
  Fix rocm cmake (#32230) · f3e49c40
  由 Qi Li 提交于 4月 14, 2021
```
* [ROCM] fix some typo in cmake, test=develop

* [ROCM] fix rccl in paddle build script, test=develop
```
  f3e49c40
- T
  Delete grpc.cmake/distribeted/distributed_ops (#32166) · 22ea4c30
  由 tianshuo78520a 提交于 4月 14, 2021
```
* Delete grpc.cmake/distribeted/distributed_ops

* reset operators/CMakeLists.txt

* rm test_transpiler_ops.py

* del test_transpiler_ops.py
```
  22ea4c30
- Z
  fix matrix_inverse_op with rocm (#32128) · 995b5f2c
  由 zhulei 提交于 4月 14, 2021
```
* fix matrix_inverse_op with rocm

* fix matrix_inverse_op with rocm

* fix matrix_inverse_op with rocm

* fix matrix_inverse_op with rocm
```
  995b5f2c
- X
  
  Add model benchmark ci (#32247) · 279b653c
  由 xiegegege 提交于 4月 14, 2021
  
  279b653c
- F
  add common dtypes as paddle's dtypes (#32012) · 95939b52
  由 Feiyu Chan 提交于 4月 14, 2021
```
* add common dtypes as paddle's dtypes

* import paddle.fluid.core_avx.VarDesc.VarType as paddle.dtype
```
  95939b52
- T
  
  fix expand op lack of float16 (#32238) · f4b2ce44
  由 Thomas Young 提交于 4月 14, 2021
  
  f4b2ce44
- X
  
  add new post-quant methods (#32208) · 4281eb49
  由 XGZhang 提交于 4月 14, 2021
  
  4281eb49
13 4月, 2021 8 次提交

extend multiclass_nms unittest timeout threshold (#32214) · cb81826a

由 Pei Yang 提交于 4月 13, 2021

* extend multiclass_nms unittest timeout threshold

* adjust timeout to 200s

* temporarily disable multiclass_nms trt op teller

cb81826a

L

upgrade to oneDNN2.2.1 (fix when prim descriptor or attr contain NaN) (#32227) · b9e543f8
由 lidanqing 提交于 4月 13, 2021

b9e543f8
Z

add statistics_UT_resource.sh for imporving UT parallel level (#32220) · 1d5d3e47
由 Zhou Wei 提交于 4月 13, 2021

1d5d3e47
Y
Fix prec on windows for long args (#32218) · 7ab47e8d
由 YUNSHEN XIE 提交于 4月 13, 2021
```
* fix error for long args

* remove unneccessary code
```
7ab47e8d

add layer.to api (#32040) · 6e946e9d

由 chentianyu03 提交于 4月 13, 2021

* add layer.to api

* add layer.to api

* add layer.to api

* add the doc for Layer.to

* add input type checking

* modify assert and import bug

* format code style

* format code style

* make place support str type

* add SetGradVarBase method to set the gradient after conversion

* modify argument palce to device

* modify argument palce to device

* modify doc of layers.to API

* add xpuplace to device argument

6e946e9d

Q

[ROCM] fix depth conv2d in rocm, test=develop (#32170) · 693c7629
由 Qi Li 提交于 4月 13, 2021

693c7629
J

optimize check_finite_and_unscale_op by fused kernel, test=develop (#31954) · fdf63b4e
由 jiangcheng 提交于 4月 13, 2021

fdf63b4e

run the sample codes added by `add_sample_code` in ops.py (#31863) · 4a09c1a1

由 Ren Wei (任卫) 提交于 4月 13, 2021

* skip paddle.Tensor.<lambda>

* some file may not exists. such as version.py, it's generated by setup.py

* debug mode

* add unittests for sampcd_processor.py

* add test cases for sampcd_processor

* add test cases for sampcd_processor

* add testcases

* add test cases

* add testcases

* add testcases

* refactor, add testcases

* add import

* all files map to pool. dont split manually

* __all__ += another list

* add testcases

* add testcases

* handle个锤子啊

* this line should not removed

https://github.com/wadefelix/Paddle/commit/882e7f7c3be6c2415f58550f82be338b84f0c0ef#diff-cb0679475bf60202fd803ae05b9146989437c3f787d1502616be6c71c69d0fb1

* print -> logger

* regulate the logging infomation

* regulate the logging infomation

* logger to file

* logger

* threads or subprocesses number config

* follow the good code style

don't touch wlist.json

* run test_sampcd_processor.py, it's a unittest for sampcd_processor.py

* update unittest for sampcd_processor.py

test=document_fix

4a09c1a1

12 4月, 2021 9 次提交

C

polish custom api content for performence (#32209) · 0624ea56
由 Chen Weihang 提交于 4月 12, 2021

0624ea56

[Rocm] fix python test of multinomial (#32158) · 4b5cb22f

由 zhulei 提交于 4月 12, 2021

* [Rocm] fix python test of multinomial

* [Rocm] fix python test of multinomial

* [Rocm] fix python test of multinomial

* [Rocm] fix python test of multinomial

4b5cb22f

Optimize the process of obtaining prec_list on windows (#32123) · 8dacfb5e

由 YUNSHEN XIE 提交于 4月 12, 2021

* test,test,notest,test=windows_ci

* test,notest,test=windows_ci

* test,notest,test=windows_ci

* test,notest,test=windows_ci

* remove test code

* delete some unnecessary logs

* fix format error

* turn on added ut check on windows

8dacfb5e

A

[CustomOp]Fix description of supporting MacOS (#32192) · bb3b7906
由 Aurelius84 提交于 4月 12, 2021

bb3b7906

[ROCM] fix some unittests (#32129) · bd2a4e23

由 ronnywang 提交于 4月 12, 2021

* [ROCM] fix test_gru_rnn_op

* [ROCM] fix test_expand_op

* [ROCM] fix test_cross_entropy_loss

* [ROCM] fix test_conv_nn_grad

* [ROCM] fix test_bilinear_tensor_product_op

* [ROCM] fix elementwise_op_function

* [ROCM] fix test_lstm_cudnn_op

* [ROCM] fix test_gpu_package_without_gpu_device

* [ROCM] fix test_gru_unit_op

* [ROCM] fix test_imperative_optimizer

* [ROCM] fix rnn

* [ROCM] fix group_norm_op

* [ROCM] fix test_pool3d_api

* [ROCM] fix test_pool3d_op

bd2a4e23

L

Optimization of bilinear backward OP CUDA kernel. (#30950) · d8afe407
由 limingshu 提交于 4月 12, 2021

d8afe407
L

follow comments to refine PR 32144 (#32174) · af374ae6
由 Leo Chen 提交于 4月 12, 2021

af374ae6
W

remove PYTHON_ABI, test=document_fix (#32190) · 80698cad
由 wuhuanzhou 提交于 4月 12, 2021

80698cad
T
fix concat_grad on kunlun (#32151) · a2387ef2
由 TTerror 提交于 4月 12, 2021
```
* fix concat_grad on kunlun

* fix concat_grad on kunlun
```
a2387ef2

10 4月, 2021 2 次提交
- A
  
  Optimize the performance of the forward of log_softmax when axis is -1 and dim <= 1024 (#31630) · f8bab5b0
  由 AshburnLee 提交于 4月 10, 2021
  
  f8bab5b0
- T
  
  Ci py3 gcc5.4 (#32045) · afa3720c
  由 tianshuo78520a 提交于 4月 10, 2021
  
  afa3720c
09 4月, 2021 3 次提交

N
make high precision for avg_pool and adaptive_avg_pool when data_type is float16 (#31887) · ec2ffb68
由 niuliling123 提交于 4月 09, 2021
```
* make high precision for avg_pool
```
ec2ffb68

[NPU] cherry-pick basic NPU components/allocator/operator/executor supports from ascendrc (#32144) · ccf5709d

由 Leo Chen 提交于 4月 09, 2021

* [feature] support npu allocator (#30840)

[feature] support npu allocator

* [feature] support npu operator (#30951)

[feature] support npu operator

* [feature] support npu allocator, part 2 (#30972)

* support npu allocator

* add npu device context

* fix some compile problem

* fix some compile problem

* add npu info

* compile ok

* fix include dir

* support naive_best_fit_allocator

* run ut ok, bug failed to exit

* call aclrtResetDevice before exit

* fix aclFinilize

* add system allocatot test

* add selected_gpus in gtest

* add tensor_test for npu

* support npu op, initial commit

* add npu stream

* add elementwise_add_op

* compile ok

* fix typo

* fix elementwise_add_op_npu_test

* support op run

* test can run but failed

* change aclopExecuteV2 to aclopCompileAndExecute

* support parsing ascend rank table file (#31000)

support parsing ascend rank table file

* Fix reshape on GE graph. (#31084)

Fix reshape on GE graph

* add npu kernel for elementwise_sub and elementwise_sub_grad (#30973)

* add npu sub op

* fix typo

* rename test

* fix bug

* fix bug

* add fp16 kernel

* fix typo

* support sub grad op

* support elementwise_sub_grad op
Co-authored-by: Nfrankwhzhang <frankwhzhang@126.com>

* Fix compilation problem (#31100)

Fix compilation problem (#31100)

* fix compile

* fix code stype

* remove const_cast

* support adding correct npu op in pybind.h (#31143)

* support adding correct npu op in pybind.h

* refine code

* [NPU] Support executor with NPU (#31057)

* [NPU] Support executor with NPU

* Fix code according to reviews

* Fix code

* Add unittest for sub op npu

* refactor npu device manager (#31154)

refactor npu device manager (#31154)

* fix selected npus

* fix compile

* fix reading flags from env

* format
Co-authored-by: Nxiayanming <41795079@qq.com>
Co-authored-by: Ngongweibao <weibao.gong@gmail.com>
Co-authored-by: Nfrankwhzhang <frankwhzhang@126.com>
Co-authored-by: Nliym27 <33742067+liym27@users.noreply.github.com>

ccf5709d

S

fix unittest timeour (#32161) · a73cb679
由 Shang Zhizhou 提交于 4月 09, 2021

a73cb679

wmsofts / Paddle 与 Fork 源项目一致

wmsofts / Paddle
与 Fork 源项目一致