提交 · 002f21851231e05e1dc9c3a79d00fe0f0b99aea8 · PaddlePaddle / Paddle

14 4月, 2023 11 次提交
- Z
  
  [AMP OP&Test] Cumprod support fp16 and bf16 (#52919) · 8a850af6
  由 Zhang Zheng 提交于 4月 14, 2023
  
  8a850af6
- C
  
  【Hackathon4 No58】logcumsum logsum (#51275) · 468869e4
  由 cyberslack_lee 提交于 4月 14, 2023
  
  468869e4
- C
  
  【Hackathon4 No58】kthvalue (#51615) · 43efb979
  由 cyberslack_lee 提交于 4月 14, 2023
  
  43efb979
- C
  【Hackathon No.62】digamma, dirichlet算子FP16/BF16单测完善 (#52604) · 7ecbcc08
  由 chenxujun 提交于 4月 14, 2023
```
* Add digamma, dirichlet tests

* Fix code
```
  7ecbcc08
- S
  【Hackathon No.55】add erf FP16 test and BF16 test (#52136) · eeb4d165
  由 superwinner1 提交于 4月 14, 2023
```
* add erf FP16 test
```
  eeb4d165
- C
  
  Add angle,bmm tests (#52630) · 6d7ee668
  由 chenxujun 提交于 4月 14, 2023
  
  6d7ee668
- U
  
  [Dcu]: Add rocsparse_spmm for dcu. (#52200) · 281ea2f4
  由 umiswing 提交于 4月 14, 2023
  
  281ea2f4
- Y
  [Zero-Dim] support 0-D tensor for... · 6f41e177
  由 YangQun 提交于 4月 14, 2023
```
[Zero-Dim] support 0-D tensor for reduce/reshape/stack/prelu/expand_v2/gaussion onednn kernels (#52185)

* support 0-D tensor for reduce/reshape/stack/prelu/expand_v2/gaussion ops

* fix gaussian random mkldnn op ut
```
  6f41e177
- G
  [phi] move sequence_pool to phi - Step 2 : sequence_pool_op (#52750) · b281b221
  由 gouzil 提交于 4月 14, 2023
```
* [phi] move sequence_pool kernel to phi

* [phi] mv sequence_pooling to phi funcs

* [phi] mv sequence_pooling_test

* [phi] RollBACK `paddle/fluid/operators/sequence_ops/sequence_pool_op.cc`

* [phi][funcs] fix mutable_data

* [phi][funcs] fix mutable_data
```
  b281b221
- S
  
  fix win cu116 compile error (#52894) · 60ba559a
  由 sneaxiy 提交于 4月 14, 2023
  
  60ba559a
- Z
  
  delete unused param from swish_grad and relu6_grad (#52805) · 54e4360a
  由 zhangyuqin1998 提交于 4月 14, 2023
  
  54e4360a
13 4月, 2023 12 次提交
- S
  【Hackathon No.55】 add channel_shuffle FP16/BF16 support and tests (#51884) · 48ccb785
  由 superwinner1 提交于 4月 13, 2023
```
* No55 add channel_shuffle FP16/BF16 support and tests
```
  48ccb785
- D
  【Hackathon No57】add_fp16_bf16_for_dot & bf16_for_cross (#52426) · 205094f0
  由 Difer 提交于 4月 13, 2023
```
* add_fp_bf_for_dot & bf_for_cross

* fix error

* fix some error

* fix some error

* change something

* fix magic number
```
  205094f0
- Z
  [AMP OP&Test] Support fp16&bf16 in reduce_max (#52862) · e0e044c0
  由 Zhang Zheng 提交于 4月 13, 2023
```
* [AMP OP&Test] Support fp16&bf16 in reduce_max
```
  e0e044c0
- L
  
  Fix the parameter check error in rmsprop_kernel_xpu. (#52866) · 9dc7e5ef
  由 Leo Guo 提交于 4月 13, 2023
  
  9dc7e5ef
- C
  
  Add pixel_shuffle pixel_unshuffle fp16/bf16 (#52582) · 2aaed989
  由 chenxujun 提交于 4月 13, 2023
  
  2aaed989
- C
  
  Add overlap_add, sign tests (#52667) · cb6de765
  由 chenxujun 提交于 4月 13, 2023
  
  cb6de765
- Z
  rename PD_REGISTER_GENERAL_KERNEL (#52759) · 3a66627e
  由 zhangyuqin1998 提交于 4月 13, 2023
```
* rename PD_REGISTER_GENERAL_KERNEL

* Update feed_op.cc

* fix

* Update strings_empty_kernel.cc
```
  3a66627e
- H
  [enforce.h Decouple logging.h] Delete glog/logging.h from enforce.h (#52651) · 5664ea26
  由 HongyuJia 提交于 4月 13, 2023
```
* [enforce.h Decouple logging.h] Delete glog/logging.h from enforce.h

* Add logging.h for profiler.cc

* Add logging.h for gloo_utils.h

* Add logging.h for addmm_kernel_impl.h

* Add logging.h for addmm_grad_kernel_impl.h

* Add logging.h for p_send_kernel.cu

* Add logging.h for determinant_grad_kernel_impl.h

* Add logging.h for p_recv_kernel.cu

* Add logging.h for elementwise_grad_base.h

* Add logging.h for transfer_layout_kernel.cc

* Add logging.h for eigvals_kernel.cc and index_select_impl.h

* Add logging.h for all files in kernel directory

* Add logging.h for xpu_info.cc

* Add logging.h for xpu
```
  5664ea26
- Z
  
  delete useless cast, elementwise_mul (#52831) · 0695fb88
  由 zhupengyang 提交于 4月 13, 2023
  
  0695fb88
- U
  
  [cutlass] Sparse conv3d backward fusion (#52361) · 0b98d1aa
  由 umiswing 提交于 4月 13, 2023
  
  0b98d1aa
- Z
  
  rename_bilinear_tensor_op (#52745) · eb93b5c9
  由 zhangyuqin1998 提交于 4月 13, 2023
  
  eb93b5c9
- C
  
  [XPU] Fix instance_norm、conv2d_xpu、inplace optimizer bugs. (#52627) · fa8abeec
  由 csy0225 提交于 4月 13, 2023
  
  fa8abeec
12 4月, 2023 4 次提交

Z
Optimize performance of unique kernel (#52736) · 8cbeefea
由 Zhang Zheng 提交于 4月 12, 2023
```
* Optimize performance of unique kernel

* fix ci
```
8cbeefea

[AMP OP&Test] add fp16/bf16 unittest for pool2d op (#52288) · f9b155f9

由 Wei Shengyu 提交于 4月 12, 2023

* add bf16 support and bf16/fp16 unittest for pool2d

* add include files

* dbg

* reformat

* reformat

* modify code according to review comment

* remove duplicate code

* remove dup code

* remove useless include

* dbg

f9b155f9

Patch del (#52754) · 189e0d44

由 wangzhen38 提交于 4月 12, 2023

* [DO NOT MERGE] adadelta lr support

* [DO NOT MERGE] gpu support

* [test] follow torch

* fix acc update order

* for ci

* [bug fix] update master para

* [bug fix] update test

* [bug fix] for ci test

* for ci

* fix xpu

* [adadelta fix] del fluid head file

* for ci

* del notes

189e0d44

[AMP OP&Test] support bf16 for batch norm (#52407) · 523f8a26

由 Guoxia Wang 提交于 4月 12, 2023

* [AMP OP&Test] support bf16 for batchnorm

* codestyle

* Update batch_norm_grad_kernel.cu

* Update batch_norm_kernel.cu

* fix codestyle

* fix

* fix

* fix

* fix

* fix

* Update batch_norm_kernel.cc

523f8a26

11 4月, 2023 7 次提交
- W
  
  [XPU] fix error pattern and rename max name (#52726) · 259b0aad
  由 wz1qqx 提交于 4月 11, 2023
  
  259b0aad
- Z
  
  delete remote_prefetch (#52748) · 3951c40d
  由 zhangyuqin1998 提交于 4月 11, 2023
  
  3951c40d
- W
  [AMP OP&Test]Add fp16/bf16 support isnan/isfinite/isinf op (#52259) · aaf873b2
  由 WJJ1995 提交于 4月 11, 2023
```
* add bfp16 test for isfinite

* fixed for ci

* deal with comments

* fixed test

* skip test in cpu

* deal with comments

* fixed for ci

* fixed testcase

* fixed for ci

* fixed for testcase
```
  aaf873b2
- W
  
  [BUG Fixs] adadelta lr support (#49732) · 23032590
  由 wangzhen38 提交于 4月 11, 2023
  
  23032590
- L
  Add output defs for eigh kernel (#51362) · da0c7e14
  由 LinearTemporalLogic 提交于 4月 11, 2023
```
* Add output defs for eigh kernel

* fix

* update

* update

* fix

* fix
```
  da0c7e14
- T
  
  [AMP OP&Test] add bf16 fp16 type support for expand_v2_op and top_k_v2_op (#51263) · 5b09dd56
  由 Thomas Young 提交于 4月 11, 2023
  
  5b09dd56
- Y
  
  update xpu.cmake to 20230408 (#52409) · 757aa470
  由 ykkk2333 提交于 4月 11, 2023
  
  757aa470
10 4月, 2023 6 次提交

D
【Hackathon No57】 add fp16 & bf16 for flip, fp16 for gaussian (#52380) · 2b0fffc2
由 Difer 提交于 4月 10, 2023
```
* add_fp_bf_for_flip_gaussian_random

* forget convert uint

* fix some error

* fix some error
```
2b0fffc2
C

【Hackathon4 No58】fix exponential and pad (#51300) · 3ee2b237
由 cyberslack_lee 提交于 4月 10, 2023

3ee2b237

[enforce.h Decouple gflags.h] Move gflags.h from enforce.h to enforce.cc (#52573) · 3c0b1795

由 HongyuJia 提交于 4月 10, 2023

* [enforce.h Decouple gflags.h] Move gflags.h from enforce.h to enforce.cc

* Add gflags.h for other files

* Add gflags.h for other files

* Add gflags.h for blas_impl.hip.h

* Add gflags.h for miopen_helper.h

3c0b1795

[AMP OP&Test] Add fp16 and bf16 test to activation (#52521) · 6bd5fd75

由 Vvsmile 提交于 4月 10, 2023

* adjust defalut tolerance of output and grad

* fix a bug in the grad of OpTest

* fix the type of setting defalut value in optest, both forward and
backward

* add defalut

* fix test_sum_op

* adjust tolerance

* fix the tolerance of eager

* add bf16 and fp16 to the activation tests

* remove some fixs

* fix activation

* fix fp16

* fix gelu

* fix the activation tests

* add bfloat16 specialization to singrad and cosgrad

* fix bugs

* fix bugs

* add unittest

* add skip

* add fp/bf to rrelu/rrelu_grad

* git add rrelu

* fix bugs

6bd5fd75

【AMP OP&Test】instance_norm fp16 and bf16 support. (#52241) · 7c98abd9

由 qizhaoaoe 提交于 4月 10, 2023

* add fp16 and bf16 support for instance_norm

* fix /= operator which not support bf16

* fix instance_norm_grad kernel and unittests.

* fix fp32 unittests.

* fix instance_norm_kernel and unittests.

* fix instance_norm_grad_kernel and unittest threshold.

* add fp16/bf16 for instance_norm_grad_grad op.

* add bf16 dtype check.

* fix conflicts.

* fix cpu support for fp32 op and fix type in instance_norm_grad_kernel.

* fix type in instance_norm_kernel.

* fix bf16 outputs in unittests and refine codes.

* fix dx computation.

* delete unuseful params and head including.

* add fp16/bf16 for static graph.

* fix device condiction for instance_norm op.

* fix instance_norm_grad_grad and bf16 op tests.

* fix op_test to support grad of bf16 can be compared with fp32.

* remove updates.

* add self-defined grad.

7c98abd9

【PaddlePaddle Hackathon 4 No.36】为 Paddle 优化 tile op 在 GPU 上的计算性能 (#52482) · 61fe2198

由 Zero Rains 提交于 4月 10, 2023

* fix divide zero bug for softmax_with_cross_entropy

* change the single test way

* can run but slow. the most important is that I do not know why it slow

* remove some useless commet

* change the copyright to correct

* remove some useless change

* if repeat_times == 1, we will not use BroadcastKernel

61fe2198

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功