提交 · 8497e2aad3dd089ea2747f2e2b159e7eb5846f71 · BaiXuePrincess / Paddle

02 3月, 2021 1 次提交

[NPU] add npu kernel for elementwise_add_grad (#31347) · 8497e2aa

由 Leo Chen 提交于 3月 02, 2021

* fix reading flags from env

* fix problem caused by async run

* support partial grad

* support elementwise_add_grad npu kernel

* add unittest

* fix bug?

8497e2aa

01 3月, 2021 3 次提交
- L
  add allreduce and broadcast without test (#31024) · 9fcdaeba
  由 lw921014 提交于 3月 01, 2021
```
add allreduce and broadcast without test
```
  9fcdaeba
- L
  
  [NPU] Support npu op: (1) slice (2) slice_grad (#31275) · a1ddff81
  由 liym27 提交于 3月 01, 2021
  
  a1ddff81
- L
  support list of list attribute for NPU (#31299) · d23bf89c
  由 Leo Chen 提交于 3月 01, 2021
```
* support list of list attribute for NPU

* fix compile problem

* fix reference
```
  d23bf89c
26 2月, 2021 1 次提交
- L
  [NPU] Support npu op pow and pow grad (#31247) · 187248f5
  由 liym27 提交于 2月 26, 2021
```
* [NPU] Support npu op: (1) pow (2) pow_grad

* Support fp16
```
  187248f5
25 2月, 2021 2 次提交
- L
  
  Fix typo of selected_npus (#31230) · d45f5d78
  由 Leo Chen 提交于 2月 25, 2021
  
  d45f5d78
- L
  refactor npu device manager (#31154) · ff4654e2
  由 Leo Chen 提交于 2月 25, 2021
```
refactor npu device manager (#31154)
```
  ff4654e2
23 2月, 2021 2 次提交
- L
  [NPU] Support executor with NPU (#31057) · 1435b4c0
  由 liym27 提交于 2月 23, 2021
```
* [NPU] Support executor with NPU

* Fix code according to reviews

* Fix code

* Add unittest for sub op npu
```
  1435b4c0
- L
  Fix compilation problem (#31100) · 85cbd556
  由 Leo Chen 提交于 2月 23, 2021
```
Fix compilation problem (#31100)
```
  85cbd556
22 2月, 2021 1 次提交

add npu kernel for elementwise_sub and elementwise_sub_grad (#30973) · 5cb20f30

由 Leo Chen 提交于 2月 22, 2021

* add npu sub op

* fix typo

* rename test

* fix bug

* fix bug

* add fp16 kernel

* fix typo

* support sub grad op

* support elementwise_sub_grad op
Co-authored-by: Nfrankwhzhang <frankwhzhang@126.com>

5cb20f30

09 2月, 2021 3 次提交

[feature] support npu allocator, part 2 (#30972) · 1201cd2e

由 Leo Chen 提交于 2月 09, 2021

* support npu allocator

* add npu device context

* fix some compile problem

* fix some compile problem

* add npu info

* compile ok

* fix include dir

* support naive_best_fit_allocator

* run ut ok, bug failed to exit

* call aclrtResetDevice before exit

* fix aclFinilize

* add system allocatot test

* add selected_gpus in gtest

* add tensor_test for npu

* support npu op, initial commit

* add npu stream

* add elementwise_add_op

* compile ok

* fix typo

* fix elementwise_add_op_npu_test

* support op run

* test can run but failed

* change aclopExecuteV2 to aclopCompileAndExecute

1201cd2e

L
[feature] support npu operator (#30951) · 7e049108
由 Leo Chen 提交于 2月 09, 2021
```
[feature] support npu operator
```
7e049108
L
[feature] support npu allocator (#30840) · 81138239
由 Leo Chen 提交于 2月 09, 2021
```
[feature] support npu allocator
```
81138239

08 2月, 2021 1 次提交
- G
  Destroy session first. (#30954) · ebef6601
  由 gongweibao 提交于 2月 08, 2021
```
Destroy session first.
```
  ebef6601
28 1月, 2021 1 次提交
- L
  Dev/fix ascend string (#30749) · 88dfd067
  由 Leo Chen 提交于 1月 28, 2021
```
Dev/fix ascend string
```
  88dfd067
27 1月, 2021 1 次提交
- L
  fix compilation on ascend-20.1 (#30722) · 6eabbc80
  由 Leo Chen 提交于 1月 27, 2021
```
fix compilation on ascend-20.1
```
  6eabbc80
21 1月, 2021 2 次提交
- G
  Add Hccl program group (#30642) · e4287ca6
  由 gongweibao 提交于 1月 21, 2021
```
Add Hccl program group
```
  e4287ca6
- G
  Add distribution supported (#30578) · f9c97dd7
  由 gongweibao 提交于 1月 21, 2021
```
Add distribution supported
```
  f9c97dd7
15 1月, 2021 5 次提交
- G
  Fix compilcation on CANN20.1 and older (#30494) · 1882f2ce
  由 gongweibao 提交于 1月 15, 2021
```
Fix compilcation on CANN20.1 and older 
```
  1882f2ce
- H
  
  Ascend rc (#30483) · 6dd52c5b
  由 hutuxian 提交于 1月 15, 2021
  
  6dd52c5b
- 石
  
  export global google flags to users, test=develop (#30448) · 715d8628
  由石晓伟提交于 1月 15, 2021
  
  715d8628
- W
  
  fix cache key for inplaced elementwise ops (#30404) · 88fc7a7d
  由 Wojciech Uss 提交于 1月 15, 2021
  
  88fc7a7d
- W
  fix the rnn mask memory bug for out of read (#30459) · 3d49882e
  由 wawltor 提交于 1月 15, 2021
```
* fix the rnn mask memory bug for out of read

* update the code for the rnn
```
  3d49882e
14 1月, 2021 5 次提交
- T
  
  support transformer v2.0 (#30381) · 6a3c8725
  由 taixiurong 提交于 1月 14, 2021
  
  6a3c8725
- S
  
  fix flatten api grad (#30426) · e85be1b1
  由 ShenLiang 提交于 1月 14, 2021
  
  e85be1b1
- Y
  
  Heter ps new (#30198) · 6e0da01c
  由 yaoxuefeng 提交于 1月 14, 2021
  
  6e0da01c
- 1
  test=develop, add distributed_infer (#30300) · 2a98e932
  由 123malin 提交于 1月 14, 2021
```
* test=develop, add distributed_infer
```
  2a98e932
- Q
  
  fix bug that cann't find mkldnn(kunlun) (#30394) · cf786d22
  由 QingshuChen 提交于 1月 14, 2021
  
  cf786d22
13 1月, 2021 8 次提交

C
skip quantizing ops in cpu inference (#30342) · 8e3a2940
由 cc 提交于 1月 13, 2021
```
* skip quantizing ops in cpu inference, test=develop
```
8e3a2940

Added support for inference using quantization aware trained dygraph (#30288) · 7bbf3ac5

由 alncat 提交于 1月 13, 2021

* added support for inference using qunatization aware trained dygraph

* added support for inference using qunatization aware trained dygraph
correct boost get usage

* Delete incorrect warning message (#30196)

* fix warning and no grad

* clean redundant API alias in 2.0 - part 2 (#30013)

* delete paddle.nn.functional.assign

* fix dynamic to static error

* just add the op error message for the matmul xpu (#30246)

 add the op error message for the matmul xpu

* Add Static Variable Clone (#30208)

Add clone method for static Variable so that this interface will be same as dygraph. It fixed some bugs in dy2stat

* use wget to replace curl to download the lcov file (#30229)

* use wget to replace curl to download the lcov file

* add cache for lcov

* fix test_pool3d_op timeout issue (#30248)

* Fix unittests bugs. (#30250)

* modify error message based on comments (#30189)

* modify error message based on comments

* edit code according to review.

* Correct spelling according to review.

* Fix bug for 'save mutiple method' (#30218)

* Fix bug for 'save mutiple method'

* To pass coverage.

* edit code to pass coverage.

* edit code to pass coverage.

* add unittest for coverage.

* change for coverage.

* edit for coverage.

* added support for inference using qunatization aware trained dygraph

* Alias from  paddle.fluid.layers.auc to paddle.static.auc (#30206)

* add alias from  fluid.layers.auc to static.auc

* Update __init__.py

* added support for inference using qunatization aware trained dygraph
correct boost get usage

* corrected boost get usage

* corrected naming issues and enforcing zero check

* correct paddle enforce message

* added more error checkings

* corrected error report message and optimized code

* corrected findvar usage

* corrected paddle_enforce in scope

* correct error messages

* correct error reporting format
Co-authored-by: NLielinJiang <50691816+LielinJiang@users.noreply.github.com>
Co-authored-by: NXiaoguangHu <46782768+XiaoguangHu01@users.noreply.github.com>
Co-authored-by: Nwawltor <fangzeyang0904@hotmail.com>
Co-authored-by: NHuihuang Zheng <zhhsplendid@gmail.com>
Co-authored-by: NYUNSHEN XIE <1084314248@qq.com>
Co-authored-by: NBai Yifan <me@ethanbai.com>
Co-authored-by: Ngongweibao <weibao.gong@gmail.com>
Co-authored-by: NWeiXin <weixin10@baidu.com>
Co-authored-by: NJiaqi Liu <liujiaqi06@baidu.com>

7bbf3ac5

G
Softmax backward optimize (#30249) · 180877e9
由 GaoWei8 提交于 1月 13, 2021
```
* softmax backward optimize
```
180877e9

fix bug on compiling inference shared lib with crypto;test=develop (#30269) · 10a8f3e5

由 Zhang Jun 提交于 1月 13, 2021

* fix bug on compiling inference shared lib with crypto;test=develop

* fix cmake bug when build inference lib using -DWITH_CRYPTO=OFF

* update cmake

* remove unnecessary enforce message

10a8f3e5

Fix Sleep Error in enforce.h (#30335) · 28e156c2

由 Huihuang Zheng 提交于 1月 13, 2021

usleep function in <unistd.h> only takes argument less than 1,000,000. Current call can exceed this limit, we have to fix it. This PR can fix random CI error.

28e156c2

Set expected place in child thread for dataloader to avoid costing cuda memory... · 3d015f1c

由 Leo Chen 提交于 1月 13, 2021

Set expected place in child thread for dataloader to avoid costing cuda memory on other card (#30338)

* set expected place in child thread for dataloader

* set device id when set tensor from numpy

* revert tensor_py change

* add compile guard

* fix ci

* fix bug

3d015f1c

Q
optimize memcpy perf for kunlun (#30291) · 2c1bba02
由 QingshuChen 提交于 1月 13, 2021
```
* optimize memcpy perf for kunlun

* remove useless unitest for kunlun mean

* minor
```
2c1bba02
S

Support unused parameters in dynamic graph distributed (#30224) · a60f17b8
由 ShenLiang 提交于 1月 13, 2021

a60f17b8

12 1月, 2021 4 次提交
- J
  
  Recompute Offload (#30233) · 75936d83
  由 JZ-LIANG 提交于 1月 12, 2021
  
  75936d83
- L
  
  correct the allowed dimension size (#30326) · a60893f6
  由 lidanqing 提交于 1月 12, 2021
  
  a60893f6
- C
  
  remove c++ stacktrace hint (#30325) · c8c8f205
  由 Chen Weihang 提交于 1月 12, 2021
  
  c8c8f205
- T
  add sparse embedding & load vars for 2.0 & gloo bug fix (#30306) · 5e839e4d
  由 tangwei12 提交于 1月 12, 2021
```
* add sparse embedding & load vars for 2.0

Change-Id: I36b59ed5f015189dc9d9d2e34a9357722d369f1b

* fix hdfs gloo

Change-Id: Ia84d579053720ad804183e54c9a04b4f031c79c6

* fix gloo hdfs

Change-Id: I5ab982fd483cddc10adcdef0b8aa83aca976cb9e

* move loadvar/sparse embedding from incubute to static

Change-Id: I57081d3545ad2efab78c72420d2162c0eacaf3a0
```
  5e839e4d

BaiXuePrincess / Paddle 与 Fork 源项目一致

BaiXuePrincess / Paddle
与 Fork 源项目一致