提交 · 39dcfc6c3a0a92f248e2ca80480df4695fdb4799 · Crayon鑫 / Paddle

14 9月, 2021 2 次提交
- P
  
  fix elementwise_div npu op (#35700) · 0bbff93c
  由 pangyoki 提交于 9月 14, 2021
  
  0bbff93c
- Y
  Implement FunctionTraits to support two kinds of elementwise functor and... · 12bf0502
  由 Yiqun Liu 提交于 9月 14, 2021
```
Implement FunctionTraits to support two kinds of elementwise functor and remove some old codes for broadcast. (#35688)
```
  12bf0502
13 9月, 2021 2 次提交
- Y
  Revert "Implement FunctionTraits to support two kinds of elementwise functor... · 40d4a295
  由 Yiqun Liu 提交于 9月 13, 2021
```
Revert "Implement FunctionTraits to support two kinds of elementwise functor and remove some old codes for broadcast. (#35487)" (#35686)
```
  40d4a295
- Y
  Implement FunctionTraits to support two kinds of elementwise functor and... · d4f84d46
  由 Yiqun Liu 提交于 9月 13, 2021
```
Implement FunctionTraits to support two kinds of elementwise functor and remove some old codes for broadcast. (#35487)
```
  d4f84d46
08 9月, 2021 1 次提交
- W
  multiply supports bool · db5fd2a1
  由 will-jl944 提交于 9月 08, 2021
```
multiply supports bool  
```
  db5fd2a1
07 9月, 2021 1 次提交
- N
  
  Modify the elementwise op according to the kernel primitive API (#34456) · eae4bf5b
  由 niuliling123 提交于 9月 07, 2021
  
  eae4bf5b
06 9月, 2021 1 次提交
- W
  Add the extra flag for the some ops (#35442) · 49797d85
  由 wawltor 提交于 9月 06, 2021
```
* Add the extra flag for the some ops

* fix the compile problem in matmul extra
```
  49797d85
03 9月, 2021 2 次提交
- Y
  
  Unify the implementation of AlignedVector and simplify the codes of dropout and cast. (#35373) · c171eca2
  由 Yiqun Liu 提交于 9月 03, 2021
  
  c171eca2
- W
  [NPU] Add elementwise_pow_grad npu op (#35278) · e913796c
  由 WJJ1995 提交于 9月 03, 2021
```
* add elementwise_pow_grad_npu

* fixed bug for CI

* deal with comments

* fixed bug for CI

* deal with comments
```
  e913796c
02 9月, 2021 1 次提交
- W
  add axis check for elementwise op while the dimension of x is equal to the... · 25871e0e
  由 wangxinxin08 提交于 9月 02, 2021
```
add axis check for elementwise op while the dimension of x is equal to the dimension of tensor (#35340)
```
  25871e0e
31 8月, 2021 1 次提交
- A
  
  NPU add elementwise_mod (#35245) · 561841d2
  由 Aganlengzi 提交于 8月 31, 2021
  
  561841d2
27 8月, 2021 1 次提交

add elementwise max grad op for npu (#34862) · 5310ceab

由 baoachun 提交于 8月 27, 2021

* add elementwise max grad op for npu

* add elementwise max grad op for npu

* add elementwise max grad op for npu

* add elementwise max grad op for npu

* add elementwise max grad op for npu

5310ceab

26 8月, 2021 1 次提交

[oneDNN] disable caching oneDNN primitives in matmul v2, Reduce grad and... · 31f0221f

由 Jacek Czaja 提交于 8月 26, 2021

[oneDNN] disable caching oneDNN primitives in  matmul v2, Reduce grad and elementwise_add grad, expand_v2 (#35132)

* - grad caching disabled of matmul_v1

- compilation fix

- compilation fix

* - reduction removed

* - Matmul v2 disabled caching

* Draft of further changes

* - workaround for reducegrad

* - fixes to UT

* - fix to compilation

* - another fix

* - fix

31f0221f

25 8月, 2021 2 次提交
- R
  
  [NPU] Fix the performance problem when 'axis' is not specified (#35116) · 91ba86b1
  由 ronnywang 提交于 8月 25, 2021
  
  91ba86b1
- T
  
  update elementwise api in kunlun (#35021) · ff96a7d5
  由 taixiurong 提交于 8月 25, 2021
  
  ff96a7d5
22 8月, 2021 1 次提交
- Z
  
  implementation of broadcast add backward by reduce (#34143) · 56c5e210
  由 Zhang Zheng 提交于 8月 22, 2021
  
  56c5e210
16 8月, 2021 1 次提交

[oneDNN] Fix to 34554 (same as previous PR but should build with GPU) (#34859) · 9cb65653

由 Jacek Czaja 提交于 8月 16, 2021

* - Added softmax without caching

* - Binary is no longer manually cached

* - Activation onednn caching removed

* - Removed manual caching of activation

* - modified UT

* - fix

* - fix

* - fixes to building

* - fix

* - fix

* - fix to UT

* - Faulty UT workaround

* - approval workaround

* - Fixes after review

* - compilation fixes

* - more lint fixes

* - more fixes after review

* - fixes after another round of review

* - hopefully compilation fix

- compilation fix

9cb65653

12 8月, 2021 1 次提交
- C
  Revert "[oneDNN] Fix to issue #34554 (#34623)" (#34838) · dc62a227
  由 Chen Weihang 提交于 8月 12, 2021
```
This reverts commit 0a5c99e8.
```
  dc62a227
11 8月, 2021 2 次提交

[oneDNN] Fix to issue #34554 (#34623) · 0a5c99e8

由 Jacek Czaja 提交于 8月 11, 2021

* - Added softmax without caching

* - Binary is no longer manually cached

* - Activation onednn caching removed

* - Removed manual caching of activation

* - modified UT

* - fix

* - fix

* - fixes to building

* - fix

* - fix

* - fix to UT

* - Faulty UT workaround

* - approval workaround

* - Fixes after review

* - compilation fixes

* - more lint fixes

* - more fixes after review

* - fixes after another round of review

0a5c99e8

A

[NPU] add elementwise_min_grad_op_npu,test=develop (#34731) · 45af4f2a
由 andyjpaddle 提交于 8月 11, 2021

45af4f2a

09 8月, 2021 1 次提交

[NPU] add broadcast supporting for elementwise_add_op_npu (#34057) · b7355d8e

由 ronnywang 提交于 8月 08, 2021

* add broadcast supporting for elementwise_add

* add broadcast supporting for elementwise_add

* add more tests

* remove the redundant code

* update

* fix place error in unittest

* remove skip.If

b7355d8e

05 8月, 2021 1 次提交
- L
  
  Support Ternary ops in elmentwise and broadcast (#33976) · 1d7b75dd
  由 limingshu 提交于 8月 05, 2021
  
  1d7b75dd
07 7月, 2021 1 次提交
- T
  
  [xpu] add dropout & amp ops in xpu place (#33891) · 84e813e3
  由 taixiurong 提交于 7月 07, 2021
  
  84e813e3
05 7月, 2021 2 次提交
- W
  
  Add fused elemwise gelu and optimize performance (#33480) · eae31856
  由 WangXi 提交于 7月 05, 2021
  
  eae31856
- L
  Enhance error message when x or y is empty in elementwise_op (#33928) · 70100e4f
  由 Leo Chen 提交于 7月 05, 2021
```
* enhance error message when x or y is empty in elementwise_op

* format code

* format code
```
  70100e4f
24 6月, 2021 1 次提交
- J
  [oneDNN] Fix to #33282 , added support of X input broadcasting to oneDNN elementwise ops (#33549) · 049dd853
  由 Jacek Czaja 提交于 6月 24, 2021
```
* - fix to #33282

* - Increased threshold for elementwise_mul_bf16 grad

* -disabled faulty UT

* - fix to approval
```
  049dd853
23 6月, 2021 1 次提交
- L
  
  Support Mod in elementwise system (#33052) · 10171806
  由 limingshu 提交于 6月 23, 2021
  
  10171806
12 6月, 2021 1 次提交
- L
  
  Support Div and FloorDiv functor in elementwise system (#33053) · fcd93b32
  由 limingshu 提交于 6月 12, 2021
  
  fcd93b32
04 6月, 2021 1 次提交
- L
  
  Reimplement logical functors with the new optimized elementwise function (#33089) · 941308c2
  由 limingshu 提交于 6月 04, 2021
  
  941308c2
02 6月, 2021 2 次提交
- L
  
  Support Add Sub Mul Max Min Pow binary functors in elementwise system (#33050) · b432d024
  由 limingshu 提交于 6月 02, 2021
  
  b432d024
- L
  
  Reimplement the comparision binary ops using the new optimized CUDA function (#33064) · 0f154961
  由 limingshu 提交于 6月 02, 2021
  
  0f154961
26 5月, 2021 1 次提交

[NPU] refine NpuOpRunner (#32869) · 8259d9bf

由 Leo Chen 提交于 5月 26, 2021

* refine ~npuOpRunner

* implement destructor and forbid copy

* use reference to avoid copy

* use const reference

* relax adam precision

* fix top_k

8259d9bf

25 5月, 2021 1 次提交

modify complex template for elementwise ops (#33071) · dbc08d69

由 chentianyu03 提交于 5月 25, 2021

* modify complex template for elementwise ops

* modify mul, div grad struct

* add complex template for CudaShuffleDownSync CudaShuffleXorSync funcs and fix the bug when delete cuda<9000

* fix shuffle func args bug

* fix shuffle func args bug

* fix shuffle func args bug

dbc08d69

24 5月, 2021 1 次提交
- L
  
  Support OutType tmeplate argument in elementwise_broadcast branch (#33060) · d6aea4ac
  由 limingshu 提交于 5月 24, 2021
  
  d6aea4ac
20 5月, 2021 2 次提交
- T
  fix gather op and add logsumexp op on kunlun (#32931) · a96e8bc9
  由 TTerror 提交于 5月 20, 2021
```
* fix gather op and add logsumexp op on kunlun

* update xpu depence

* update tests and fix elementwise_add
```
  a96e8bc9
- L
  
  Binary functor envoking of elementwise broadcast (#32928) · 14949521
  由 limingshu 提交于 5月 20, 2021
  
  14949521
14 5月, 2021 1 次提交
- L
  
  Optimization the broadcast performance of elementwise_add (#32512) · b035c8b0
  由 limingshu 提交于 5月 14, 2021
  
  b035c8b0
10 5月, 2021 1 次提交
- Z
  
  Support different data type between input and output (#32823) · 3419de53
  由 Zhang Zheng 提交于 5月 10, 2021
  
  3419de53
22 4月, 2021 1 次提交
- Z
  
  Modify some contents for elementwise op impl (#32414) · 890d6bc0
  由 Zhang Zheng 提交于 4月 22, 2021
  
  890d6bc0
19 4月, 2021 1 次提交

[NPU] cherry-pick gc/dataloader/save&load/optimization from ascendrc to develop (#32294) · cbe5c9f8

由 Leo Chen 提交于 4月 19, 2021

* [NPU] support GarbageCollector for npu (#31874)

* support GarbageCollector for npu

* fix typo

* fix gather_grad

* disable NPUDefaultStreamGarbageCollector on NPU

* [NPU] support npu for memcpy op (#31808)

* support npu for memcpy op

* add ut

* fix ut

* fix typo

* 【NPU】fix bug of using temp vector (#31963)

* fix bug when beta1_pow on cpu (#31995)

* [NPU] support npu profiler (#31684)

* support npu profiler

* add python api

* fix bugs

* add wrapper for incomplete type

* update profile proto

* record npu wait

* add xpu placeholder

* fix adam (#32016)

* [NPU] enable async copy and  add wait before sync operation (#31956)

* enable async copy and  add wait before sync operation

* remove unneccessary wait

* add FillNpuTensorWithConstant

* refine

* fix fill_constant

* make TensorFromVector/TensorToVector sync

* [NPU] Support dataloader on npu place. (#31867)

* [NPU] Wait on NPUPlace (#32086)

* [NPU] fix cast op (#32121)

* fix npu kernel of cast op to handle casting to same dtype

* add comments

* [NPU] support cann 20.3 (#32044)

* fix compile problem on cann 20.3

* fix ut

* fix test_mul

* fix check_finite_and_scale

* fix lookup_table_v2_grad

* fix cmake

* support print op

* [NPU] Support npu save load (#31893)

* support save load for NPU

* add save load npu unittest

* support np.array transform in NPU

* fix errors

* delete dygraph in unittest

* add Wait

* fix unittest

* fix review comment

* fix unittest problem

* fix little problem

* change aclrtSynchronizeDevice to aclrtSynchronizeStream for better performance (#32196)

* change aclrtSynchronizeDevice to aclrtSynchronizeStream for better performace

* refine code

* fix NPUDeviceContext in all c++ unittest (#32198)

* fix NPUDeviceContext in all c++ unittest

* refine log
Co-authored-by: Npangyoki <pangyoki@126.com>

* [NPU] Remove TensorFromVector and avoid sync copy in npu op kernel for better performance (#31994)

* enable async copy and  add wait before sync operation

* remove unneccessary wait

* add FillNpuTensorWithConstant

* refine

* fix fill_constant

* change TensorFromVector to FillNpuTensorWithConstant

* fix ignored api

* delete extra unittest

* fix little error

* fix update_loss_scaling_op_npu and check_finite_and_unscale_op_npu

* change TensorCopySync to TensorCopy

* delete useless Wait and add StreamWait

* fix npu_stream error

* fix check_finite_and_unscale_op_npu TensorCopy

* only save stream wait

* fix NPUDeviceContext in all c++ unittest

* delete wait
Co-authored-by: Nzhiqiu <chenqiuliang@baidu.com>

* delete useless unittest file (#32206)

* Fix op test (#32231)

* fix conditional block (#32243)

* fix adam bug again (#32246)

* fix compile

* fix ut

* fix ut
Co-authored-by: Nliym27 <33742067+liym27@users.noreply.github.com>
Co-authored-by: Npangyoki <pangyoki@126.com>

cbe5c9f8

Crayon鑫 / Paddle 与 Fork 源项目一致

Crayon鑫 / Paddle
与 Fork 源项目一致