提交 · 995b5f2c0f0135ea00119f4a247c9bd7385913c0 · 机器未来 / Paddle

14 4月, 2021 1 次提交

fix matrix_inverse_op with rocm (#32128) · 995b5f2c

由 zhulei 提交于 4月 14, 2021

* fix matrix_inverse_op with rocm

* fix matrix_inverse_op with rocm

* fix matrix_inverse_op with rocm

* fix matrix_inverse_op with rocm

995b5f2c

13 4月, 2021 2 次提交
- Q
  
  [ROCM] fix depth conv2d in rocm, test=develop (#32170) · 693c7629
  由 Qi Li 提交于 4月 13, 2021
  
  693c7629
- J
  
  optimize check_finite_and_unscale_op by fused kernel, test=develop (#31954) · fdf63b4e
  由 jiangcheng 提交于 4月 13, 2021
  
  fdf63b4e
12 4月, 2021 3 次提交

[ROCM] fix some unittests (#32129) · bd2a4e23

由 ronnywang 提交于 4月 12, 2021

* [ROCM] fix test_gru_rnn_op

* [ROCM] fix test_expand_op

* [ROCM] fix test_cross_entropy_loss

* [ROCM] fix test_conv_nn_grad

* [ROCM] fix test_bilinear_tensor_product_op

* [ROCM] fix elementwise_op_function

* [ROCM] fix test_lstm_cudnn_op

* [ROCM] fix test_gpu_package_without_gpu_device

* [ROCM] fix test_gru_unit_op

* [ROCM] fix test_imperative_optimizer

* [ROCM] fix rnn

* [ROCM] fix group_norm_op

* [ROCM] fix test_pool3d_api

* [ROCM] fix test_pool3d_op

bd2a4e23

L

Optimization of bilinear backward OP CUDA kernel. (#30950) · d8afe407
由 limingshu 提交于 4月 12, 2021

d8afe407
T
fix concat_grad on kunlun (#32151) · a2387ef2
由 TTerror 提交于 4月 12, 2021
```
* fix concat_grad on kunlun

* fix concat_grad on kunlun
```
a2387ef2

10 4月, 2021 1 次提交
- A
  
  Optimize the performance of the forward of log_softmax when axis is -1 and dim <= 1024 (#31630) · f8bab5b0
  由 AshburnLee 提交于 4月 10, 2021
  
  f8bab5b0
09 4月, 2021 3 次提交

N
make high precision for avg_pool and adaptive_avg_pool when data_type is float16 (#31887) · ec2ffb68
由 niuliling123 提交于 4月 09, 2021
```
* make high precision for avg_pool
```
ec2ffb68

[NPU] cherry-pick basic NPU components/allocator/operator/executor supports from ascendrc (#32144) · ccf5709d

由 Leo Chen 提交于 4月 09, 2021

* [feature] support npu allocator (#30840)

[feature] support npu allocator

* [feature] support npu operator (#30951)

[feature] support npu operator

* [feature] support npu allocator, part 2 (#30972)

* support npu allocator

* add npu device context

* fix some compile problem

* fix some compile problem

* add npu info

* compile ok

* fix include dir

* support naive_best_fit_allocator

* run ut ok, bug failed to exit

* call aclrtResetDevice before exit

* fix aclFinilize

* add system allocatot test

* add selected_gpus in gtest

* add tensor_test for npu

* support npu op, initial commit

* add npu stream

* add elementwise_add_op

* compile ok

* fix typo

* fix elementwise_add_op_npu_test

* support op run

* test can run but failed

* change aclopExecuteV2 to aclopCompileAndExecute

* support parsing ascend rank table file (#31000)

support parsing ascend rank table file

* Fix reshape on GE graph. (#31084)

Fix reshape on GE graph

* add npu kernel for elementwise_sub and elementwise_sub_grad (#30973)

* add npu sub op

* fix typo

* rename test

* fix bug

* fix bug

* add fp16 kernel

* fix typo

* support sub grad op

* support elementwise_sub_grad op
Co-authored-by: Nfrankwhzhang <frankwhzhang@126.com>

* Fix compilation problem (#31100)

Fix compilation problem (#31100)

* fix compile

* fix code stype

* remove const_cast

* support adding correct npu op in pybind.h (#31143)

* support adding correct npu op in pybind.h

* refine code

* [NPU] Support executor with NPU (#31057)

* [NPU] Support executor with NPU

* Fix code according to reviews

* Fix code

* Add unittest for sub op npu

* refactor npu device manager (#31154)

refactor npu device manager (#31154)

* fix selected npus

* fix compile

* fix reading flags from env

* format
Co-authored-by: Nxiayanming <41795079@qq.com>
Co-authored-by: Ngongweibao <weibao.gong@gmail.com>
Co-authored-by: Nfrankwhzhang <frankwhzhang@126.com>
Co-authored-by: Nliym27 <33742067+liym27@users.noreply.github.com>

ccf5709d

Y

Advoid CPU -> CPU memory copy when start, end, step is already on CPU. (#29088) · 95122ebe
由 Yiqun Liu 提交于 4月 09, 2021

95122ebe

08 4月, 2021 2 次提交
- C
  Support converting the model from fp32 to fp16 (#32112) · 1bae1e74
  由 cc 提交于 4月 08, 2021
```
* Support converting the model from fp32 to fp16
```
  1bae1e74
- T
  
  fix the XXX_GRAD_CASE bug by HexToString (#32004) · f74f9762
  由 Thomas Young 提交于 4月 08, 2021
  
  f74f9762
07 4月, 2021 4 次提交

D
add uint8 type for flatten op (#32120) · 297290a8
由 danleifeng 提交于 4月 07, 2021
```
* add uint8 type for flatten;test=develop
```
297290a8

【NPU】Merge ascend GE&distributed code by 0208 from ascendrc (#31957) · 8c7c53b3

由 zhang wenhui 提交于 4月 07, 2021

* Ascend rc (#30483)

* Fix compilcation on CANN20.1 and older (#30494)

Fix compilcation on CANN20.1 and older

* Add distribution supported (#30578)

Add distribution supported

* Build praser for Hcom* operators (#30627)

Build praser for Hcom* operators

* Pass device_ids info from launch to trainer. (#30632)

Pass device_ids info from launch to trainer

* Add Hccl program group (#30642)

Add Hccl program group

* Add startup bash files of test_ascend_group. (#30645)

Add startup bash files of test_ascend_group

* cleanup (#30646)

cleanup test_ascend_group.py

* [Feature] Build parser to support distributed training (#30658)

[Feature] Build parser to support distributed training

* fix compilation on ascend-20.1 (#30722)

fix compilation on ascend-20.1

* Dev/fix ascend string (#30749)

Dev/fix ascend string

* code style (#30781)

code style

* Merge ascend_optimizer and ascend_parser. (#30776)

Merge ascend_optimizer and ascend_parser.

* Ascendrc add converted op : [range/equal/range/uniform_random/expand/squeeze], fix cast op bug  (#30797)

Ascendrc add converted op : [range/equal/range/uniform_random/expand/squeeze], fix cast op bug

* Add paddle ascend distribution training supported (#30796)

Add paddle ascend distribution training supported

* pass cxx_flags to gloo cmake (#30857)

* Destroy session first. (#30954)

Destroy session first.

* merge

* fix, test=develop

* fix, test=develop

* fix style, test=develop

* fix, test=develop

* fix

* fix log fatal, test=develop

* fix enforce style, test=develop

* fix, test=develop

* fix, test=develop

* fix rccl, test=develop

* fix test, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix node_num, test=develop

* fix ids str, test=develop

* fix ids str, test=develop

* fix ids str, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix, test=develop

* fix style code, test=develop

* fix style code, test=develop

* fix style code, test=develop

* fix style code, test=develop
Co-authored-by: Nhutuxian <hutuxian2011@sina.cn>
Co-authored-by: Ngongweibao <weibao.gong@gmail.com>
Co-authored-by: NVoid Main <voidmain1313113@gmail.com>
Co-authored-by: NLeo Chen <chenqiuliang@baidu.com>
Co-authored-by: Ndingsiyu <18369187719@163.com>
Co-authored-by: NOleNet <olenet@126.com>

8c7c53b3

O
improve performance of DepthwiseConv(NHWC) (#31677) · 363b25aa
由 Ouyang Chao 提交于 4月 07, 2021
```
* improve performance of DepthwiseConv(NWHC)
```
363b25aa

Struct SparseValue && Bug Fix (#31721) · a881b4d5

由 tangwei12 提交于 4月 07, 2021

* add PullSparseValue for pull sparse

* fix bug for PullSparseValue

* add test mode in lookuptable

* revert API change

* add comment for is_training

a881b4d5

06 4月, 2021 3 次提交
- W
  
  optimize compilation of operators using eigen (#31851) · 187bf412
  由 wuhuanzhou 提交于 4月 06, 2021
  
  187bf412
- K
  fix two error message (#32039) · 9e8f9037
  由 Kqnonrime 提交于 4月 06, 2021
```
* fix two error message

* fix two error message

* fix error

* fix error

* fix error

* fix error
```
  9e8f9037
- R
  
  [ROCM] fix the backward maxpool (#32030) · a3b08bad
  由 ronnywang 提交于 4月 06, 2021
  
  a3b08bad
03 4月, 2021 1 次提交
- J
  
  Optimize elementwise_add_grad op, test=develop (#32051) · 1e52f324
  由 jiangcheng 提交于 4月 03, 2021
  
  1e52f324
02 4月, 2021 2 次提交
- R
  
  [ROCM] fix softmax_with_cross_entropy_op (#31982) · 9e06a641
  由 ronnywang 提交于 4月 02, 2021
  
  9e06a641
- N
  add leaky_relu forward and backward in activation_op.cu (#31841) · 4490e8af
  由 niuliling123 提交于 4月 02, 2021
```
* add leaky_relu forward and backward in activation_op.cu
```
  4490e8af
01 4月, 2021 6 次提交

Q

[ROCM] fix depthwise conv failure on ROCM, test=develop (#31998) · a4b30a12
由 Qi Li 提交于 4月 01, 2021

a4b30a12
H

remove useless code (#32001) · 9c5d0286
由 hutuxian 提交于 4月 01, 2021

9c5d0286

[Paddle-TRT] add anchor generator op plugin (#31730) · b807e408

由 zlsh80826 提交于 4月 01, 2021

* add anchor generator op plugin

* add anchor generator unit_test

* remove dbg info

* remove redundant line

* replace assertion with paddle enforce

* dynamic plugin replaces assertion with paddle enforce

* anchor generator support dynamic shape on spatial axis

* anchor generator test with fp16, dynamic shape

* add anchor generator test all

* add back main

* reduce test input size to not exceed the timelimit of ci

* change super to InferencePassTest for python2 compatibility

* reuse paddle operator anchor generator

* move creator construct to header with default

* add cuda ifdef

* reduce line

* change super to InferencePassTest for python2 compatibility

* fix anchor generator fp16 serialize setting

* split unittest from test_all

* restrict anchor generator input format before version 7234

* anchor generator only support greater than trt7.1

* change min_graph_size to 2

* min_graph size to 3 if dynamic shape

* reduce dynamic shape size to avoid trt search tactic too long to exceed time limit

* remove anchor from fetch list

* anchor generator support all trt version

* fix memory not allocated but if serialized

b807e408

Z

Optimize the perf of SameDimsAdd CUDA Kernel (#31872) · 4acc87be
由 Zhang Zheng 提交于 4月 01, 2021

4acc87be
Z

Support uint8_t for fill_constant_op (#31911) · 980227f9
由 Zhang Zheng 提交于 4月 01, 2021

980227f9
K
new group (#31682) · 07741593
由 kuizhiqing 提交于 4月 01, 2021
```
* new group

* ci compatible fix

* assert nccl
```
07741593

31 3月, 2021 7 次提交

fix one error massage (#31904) · 6f85e241

由 Kqnonrime 提交于 3月 31, 2021

* fix one error massage

* fix a error message

* new fix three error messages

* new fix three error messages

* new fix some error

* new fix one error message

6f85e241

T

delete cuda9 code (#31883) · ea738dda
由 tianshuo78520a 提交于 3月 31, 2021

ea738dda

Update eigen version to f612df27 (#31832) · 495e7f9c

由 wuhuanzhou 提交于 3月 31, 2021

* update eigen version to f612df27, test=develop

* fix compilation error, test=develop

* remove patch command in eigen, test=develop

* fix compilation error caused by call Eigen function with float16 and bfloat16, test=develop

* fix unittest error, test=develop

* fix unittest error caused by precision, test=develop

* remove patch files used by old version eigen, test=develop

495e7f9c

update compilation with C++14 (#31815) · 587d99ae

由 wuhuanzhou 提交于 3月 31, 2021

* update compilation with C++14, test=develop

* fix compilation error in eigen, test=develop

587d99ae

T
fix split core (#31892) · 393b3bd6
由 Thunderbrook 提交于 3月 31, 2021
```
* fix split core

* format
```
393b3bd6
T

fix some bug in transformer training in xpu (#31918) · 52b05bac
由 taixiurong 提交于 3月 31, 2021

52b05bac

[ROCM] Add ROCm support for warpctc op (#31817) · ef8323d4

由 furnace 提交于 3月 31, 2021

* bugfix for warpctc

* fix warpctc commit id

* fix warpctc commit id

* fix warpctc commit id

* fix warpctc commit id

* fix warpctc commit id

* fix WARPCTC_WITH_HIP invalid

* Add logs to find out why can not dlopen libwarpctc.so

* fix warpctc commit id

* fix unit test test_warpctc_op

* Optime failed log for dlopen

* Optime failed log for dlopen

* Delete extra changes

* fix warpctc commit id

* fix warpctc commit id

* Add is_compiled_with_rocm for test_warpctc_op

* fix warpctc commit id

* Cancel optimize dlopen failed reason, move to next pr, due to it makes windows ci failed

* Cancel optimize dlopen failed reason, move to next pr, due to it makes windows ci failed

* Cancel optimize dlopen failed reason, move to next pr, due to it makes windows ci failed

* fix code style problems

ef8323d4

30 3月, 2021 2 次提交
- J
  
  fix stack op grad nullptr (#31962) · 95f808c8
  由 Jiawei Wang 提交于 3月 30, 2021
  
  95f808c8
- J
  
  Added int8 kernel for oneDNN LSTM op (#31894) · 6dca7a1d
  由 jakpiase 提交于 3月 30, 2021
  
  6dca7a1d
29 3月, 2021 3 次提交
- N
  
  relu forward and backward with vectortype (#31869) · a71d72d9
  由 niuliling123 提交于 3月 29, 2021
  
  a71d72d9
- T
  
  Delete cudnn6 code (#31835) · 8829a309
  由 tianshuo78520a 提交于 3月 29, 2021
  
  8829a309
- L
  
  Fix bug of set_value op：Decerease axes to do right broadcast (#31875) · 525c32e3
  由 liym27 提交于 3月 29, 2021
  
  525c32e3

机器未来 / Paddle 与 Fork 源项目一致

机器未来 / Paddle
与 Fork 源项目一致