提交 · 5618f14047250f1325e7d544b4c147bf0a98c5a8 · 机器未来 / Paddle

01 3月, 2021 2 次提交
- L
  
  [NPU] Support npu op: (1) slice (2) slice_grad (#31275) · a1ddff81
  由 liym27 提交于 3月 01, 2021
  
  a1ddff81
- L
  support list of list attribute for NPU (#31299) · d23bf89c
  由 Leo Chen 提交于 3月 01, 2021
```
* support list of list attribute for NPU

* fix compile problem

* fix reference
```
  d23bf89c
26 2月, 2021 1 次提交
- L
  [NPU] Support npu op pow and pow grad (#31247) · 187248f5
  由 liym27 提交于 2月 26, 2021
```
* [NPU] Support npu op: (1) pow (2) pow_grad

* Support fp16
```
  187248f5
23 2月, 2021 1 次提交
- L
  Fix compilation problem (#31100) · 85cbd556
  由 Leo Chen 提交于 2月 23, 2021
```
Fix compilation problem (#31100)
```
  85cbd556
22 2月, 2021 1 次提交

add npu kernel for elementwise_sub and elementwise_sub_grad (#30973) · 5cb20f30

由 Leo Chen 提交于 2月 22, 2021

* add npu sub op

* fix typo

* rename test

* fix bug

* fix bug

* add fp16 kernel

* fix typo

* support sub grad op

* support elementwise_sub_grad op
Co-authored-by: Nfrankwhzhang <frankwhzhang@126.com>

5cb20f30

09 2月, 2021 3 次提交

[feature] support npu allocator, part 2 (#30972) · 1201cd2e

由 Leo Chen 提交于 2月 09, 2021

* support npu allocator

* add npu device context

* fix some compile problem

* fix some compile problem

* add npu info

* compile ok

* fix include dir

* support naive_best_fit_allocator

* run ut ok, bug failed to exit

* call aclrtResetDevice before exit

* fix aclFinilize

* add system allocatot test

* add selected_gpus in gtest

* add tensor_test for npu

* support npu op, initial commit

* add npu stream

* add elementwise_add_op

* compile ok

* fix typo

* fix elementwise_add_op_npu_test

* support op run

* test can run but failed

* change aclopExecuteV2 to aclopCompileAndExecute

1201cd2e

L
[feature] support npu operator (#30951) · 7e049108
由 Leo Chen 提交于 2月 09, 2021
```
[feature] support npu operator
```
7e049108
L
[feature] support npu allocator (#30840) · 81138239
由 Leo Chen 提交于 2月 09, 2021
```
[feature] support npu allocator
```
81138239

21 1月, 2021 2 次提交
- G
  Add Hccl program group (#30642) · e4287ca6
  由 gongweibao 提交于 1月 21, 2021
```
Add Hccl program group
```
  e4287ca6
- G
  Add distribution supported (#30578) · f9c97dd7
  由 gongweibao 提交于 1月 21, 2021
```
Add distribution supported
```
  f9c97dd7
15 1月, 2021 4 次提交
- H
  
  Ascend rc (#30483) · 6dd52c5b
  由 hutuxian 提交于 1月 15, 2021
  
  6dd52c5b
- 石
  
  export global google flags to users, test=develop (#30448) · 715d8628
  由石晓伟提交于 1月 15, 2021
  
  715d8628
- W
  
  fix cache key for inplaced elementwise ops (#30404) · 88fc7a7d
  由 Wojciech Uss 提交于 1月 15, 2021
  
  88fc7a7d
- W
  fix the rnn mask memory bug for out of read (#30459) · 3d49882e
  由 wawltor 提交于 1月 15, 2021
```
* fix the rnn mask memory bug for out of read

* update the code for the rnn
```
  3d49882e
14 1月, 2021 2 次提交
- T
  
  support transformer v2.0 (#30381) · 6a3c8725
  由 taixiurong 提交于 1月 14, 2021
  
  6a3c8725
- S
  
  fix flatten api grad (#30426) · e85be1b1
  由 ShenLiang 提交于 1月 14, 2021
  
  e85be1b1
13 1月, 2021 1 次提交
- G
  Softmax backward optimize (#30249) · 180877e9
  由 GaoWei8 提交于 1月 13, 2021
```
* softmax backward optimize
```
  180877e9
12 1月, 2021 8 次提交
- J
  
  Recompute Offload (#30233) · 75936d83
  由 JZ-LIANG 提交于 1月 12, 2021
  
  75936d83
- L
  
  correct the allowed dimension size (#30326) · a60893f6
  由 lidanqing 提交于 1月 12, 2021
  
  a60893f6
- T
  Fix/distributed proto (#29981) · 25f80fd3
  由 tangwei12 提交于 1月 12, 2021
```
* rename sendrecv.proto to namespace paddle.distributed

* split ps with distributed
```
  25f80fd3
- D
  fix elugradgrad test fail & error message opt (#30171) · 231501fe
  由 Double_V 提交于 1月 12, 2021
```
* fix elugradgrad test fail and error message opt

* fix unitest,test=develop

* Update prroi_pool_op.h

fix error message

* opt message,test=develop

* fix ci fail,test=develop
```
  231501fe
- Z
  Fix the accuracy problem of allclose op when using float64 data type in static mode. (#29890) · fb49ea38
  由 Zhen Wang 提交于 1月 12, 2021
```
* Fix the accuracy problem of allclose op when using float64 data type in static mode.

* Format the code style.
```
  fb49ea38
- Y
  
  fix datanorm error msg (#30294) · 4656525e
  由 yaoxuefeng 提交于 1月 12, 2021
  
  4656525e
- F
  
  add fp16 support for tril_triu op (#30186) · 77051cc9
  由 furnace 提交于 1月 12, 2021
  
  77051cc9
- 石
  
  fix header file paths of gflags, commit 3, test=develop (#30273) · efa54629
  由石晓伟提交于 1月 12, 2021
  
  efa54629
11 1月, 2021 9 次提交
- C
  Fix server.h include device_context (#30243) · 5b2c15af
  由 Chengmo 提交于 1月 11, 2021
```
* fix cmake
Co-authored-by: NseiriosPlus <tangwei12@baidu.com>
```
  5b2c15af
- 石
  
  enhance error msgs of fusion_seqpool_cvm_concat_op.cc, test=develop (#30240) · a0ee0914
  由石晓伟提交于 1月 11, 2021
  
  a0ee0914
- L
  Support vector<double> as type of op attribute and op set_value suppport... · b4989fb7
  由 liym27 提交于 1月 11, 2021
```
Support vector<double> as type of op attribute and op set_value suppport vector<double> as value (#30126)
```
  b4989fb7
- W
  
  register OPMaker and Infer Shape Check for fused_elementwise_add (#30259) · 8dcae0c5
  由 wangchaochaohu 提交于 1月 11, 2021
  
  8dcae0c5
- A
  
  Add tf32 switch for cuDNN (#29192) · 924aac22
  由 AshburnLee 提交于 1月 11, 2021
  
  924aac22
- C
  type promotion for grad (#30177) · c7371b7b
  由 chentianyu03 提交于 1月 11, 2021
```
* type promotion for grad

* add type promotion for div op
```
  c7371b7b
- L
  
  Check the rank of input in kernel of set_value op (#30147) · 3ce878f3
  由 liym27 提交于 1月 11, 2021
  
  3ce878f3
- W
  modify error message based on comments (#30189) · 66dc4ac7
  由 WeiXin 提交于 1月 11, 2021
```
* modify error message based on comments

* edit code according to review.

* Correct spelling according to review.
```
  66dc4ac7
- W
  just add the op error message for the matmul xpu (#30246) · fee42441
  由 wawltor 提交于 1月 11, 2021
```
 add the op error message for the matmul xpu 
```
  fee42441
10 1月, 2021 2 次提交
- G
  optimize softmax forward (#30217) · 0a21924a
  由 GaoWei8 提交于 1月 10, 2021
```
* optimize softmax forward
```
  0a21924a
- W
  reduce the occupied size of memory for the fused pattern of elementwise_add... · af80859d
  由 wangchaochaohu 提交于 1月 10, 2021
```
reduce the  occupied size  of memory for the fused pattern of elementwise_add Op and activation Op(relu Op for example) (#29885)
```
  af80859d
09 1月, 2021 2 次提交
- Z
  
  enhance error message, test=develop (#30220) · 5932fee6
  由 zhang wenhui 提交于 1月 09, 2021
  
  5932fee6
- J
  [oneDNN] Added UT for testing elementwise_mul caching (#30203) · 4aba17b5
  由 Jacek Czaja 提交于 1月 09, 2021
```
* - Added UT for testing elementwise_mul caching

* lint fixes
```
  4aba17b5
08 1月, 2021 2 次提交

Support pure fp16 training for AMP API. (#29544) · 7f7dfccf

由 Zhen Wang 提交于 1月 08, 2021

* add cast ops before and after unsupported fp16 ops.

* Keep partial net in FP32 pattern.

* Support check_finite_and_unscale and update_loss_scaling for FP16 calculation mode.

* Add fp16 support for adam op.

* add multi precision attr for adam.

* Fix the bug of test_multi_precision_fp16_train UT.

* Code format for CI.

* Fix the redefine error about MPTypeTrait on windows.

* fix bugs of the _create_accumulators func in Momentum.

* fix bug when inserting post cast op.

* Add the update_loss_scaling op in allow_set of UnusedVarCheck.

* Update for ci coverage.

* Add some doc for OptimizerWithMixedPrecision.

* Fix the code style.

* Imporve the doc of `amp_init`.

* Change for fp16 testing if users have the infer program defined in separate way.

7f7dfccf

L

use cuda generator in bernoulli cuda kernel (#30199) · 789743e1
由 Leo Chen 提交于 1月 08, 2021

789743e1

机器未来 / Paddle 与 Fork 源项目一致

机器未来 / Paddle
与 Fork 源项目一致