提交 · 14a6e67b60f65539a38deb6c6a2984955d1b4d13 · PaddlePaddle / Paddle

18 11月, 2022 8 次提交
- T
  CUDNN v8 Implementation of Convolution Kernels (#47454) · 14a6e67b
  由 Tian Zheng 提交于 11月 18, 2022
```
* Refactor conv_kernel and conv_grad_kernel to provide interface for CUDNNv8 implementation

* Fix macro

* Add implementation for conv_kernel and conv_grad_kernel

* Modification after rebase onto latest develop

* Modify plan cache to comply with the API of phi::autotune

* Refactor to reduce duplicate code

* Review fix:
- move functions in  conv_kernel_impl_v8.h and conv_grad_kernel_impl_v8.h to conv_kernel.cu and conv_grad_kernelk.cu
- add const specifier for input tensor
- add logging when plans fail to execute
- move CudnnConvBwdFilterV8 and CudnnConvBwdDataV8 to conv_cudnn_frontend.h

* - move plan building outside of cache

* Fix ROCM build
```
  14a6e67b
- Y
  
  add bf16 for numel (#48121) · a7d306af
  由 Yuang Liu 提交于 11月 18, 2022
  
  a7d306af
- W
  [PHI decoupling] remove "gpu_primitives.h" in fluid (#48063) · 9918bf9c
  由 Wang Xin 提交于 11月 18, 2022
```
* remove "gpu_primitives.h" in fluid namespace

* fix PR-CI-GpuPS fail

* fix PR-CI-GpuPS fail
```
  9918bf9c
- Z
  
  cast and gradient_accumulator support double for xpu, test=kunlun (#47800) · 982d5ff7
  由 zhangyikun02 提交于 11月 18, 2022
  
  982d5ff7
- F
  
  optimize: vectorize transpose_padding (#48116) · 635958d9
  由 feng_shuai 提交于 11月 18, 2022
  
  635958d9
- F
  
  fix: supoort huge length of attention (#48053) · 42f35841
  由 feng_shuai 提交于 11月 18, 2022
  
  42f35841
- S
  
  fix onednn prelu header (#48064) · 85598e31
  由 Sylwester Fraczek 提交于 11月 18, 2022
  
  85598e31
- H
  
  rm "paddle/fluid/operators/amp/fp16_type_traits.h" in phi (#48051) · e4670d80
  由 huangjiyi 提交于 11月 18, 2022
  
  e4670d80
17 11月, 2022 18 次提交
- Z
  Clip intermediate output of op when save inference model (#48026) · fafc7be2
  由 zyfncg 提交于 11月 17, 2022
```
* clip extra and intermediate output of op

* fix bug

* fix bug

* polich code

* polich log
```
  fafc7be2
- Q
  [NPU] add _npu_identity op and api, test=develop (#47850) · 099c2302
  由 Qi Li 提交于 11月 17, 2022
```
* [NPU] add _npu_identity op and api, test=develop

* fix doc

* address comments
```
  099c2302
- W
  
  Refactor collective communication all_to_all, all_to_all_single C++ API (#48059) · 3f480af2
  由 Wen Sun 提交于 11月 17, 2022
  
  3f480af2
- W
  support int input for scale (#48044) · dbc63555
  由 wenbin 提交于 11月 17, 2022
```
* int scale

* round

* revert commit
```
  dbc63555
- X
  
  fix the thread number to ensure deterministic of embedding kernel (#48073) · 5329187d
  由 xiongkun 提交于 11月 17, 2022
  
  5329187d
- H
  
  fix new executor gc dep bug (#48068) · 04dcb9d7
  由 hong 提交于 11月 17, 2022
  
  04dcb9d7
- H
  
  rm "paddle/fluid/framework/convert_utils.h" in phi (#48001) · 2f34fc7a
  由 huangjiyi 提交于 11月 17, 2022
  
  2f34fc7a
- Y
  [PHI]Standardise some C++ API (Part5) (#47860) · f3650201
  由 YuanRisheng 提交于 11月 17, 2022
```
* standard api

* fix xpu bugs
```
  f3650201
- M
  
  optimizing a bit tensor_array initialization (#48066) · c374894d
  由 Mountagha 提交于 11月 17, 2022
  
  c374894d
- T
  
  xpu-paddlepaddle-41 [任务] ffn and attention test=kunlun (#46658) · 071708fa
  由 taixiurong 提交于 11月 17, 2022
  
  071708fa
- W
  
  move "function_traits.h" from fluid to phi (#48065) · b7841a2b
  由 Wang Xin 提交于 11月 17, 2022
  
  b7841a2b
- X
  [Paddle Inference] Support cast trt converter of bool input and output . (#48043) · ff44df18
  由 xiaoxiaohehe001 提交于 11月 17, 2022
```
* add_cast_bool

* cast
```
  ff44df18
- Y
  Implement a common dimension simplifier. (#47981) · bf6af816
  由 Yiqun Liu 提交于 11月 17, 2022
```
* Implement a common dims simplifier.

* Fix the include position error.

* Reduce the cpu overhead of broadcast computing.
```
  bf6af816
- H
  
  rm "paddle/phi/kernels/gpu/batch_norm_utils.h" in phi (#48057) · b7e120d2
  由 huangjiyi 提交于 11月 17, 2022
  
  b7e120d2
- H
  [PHI decoupling] move "paddle/fluid/operators/math.h" to phi (#48062) · f62bd3b4
  由 huangjiyi 提交于 11月 17, 2022
```
* rm "paddle/fluid/operators/math.h" in phi

* rm "paddle/fluid/operators/math.h" in fluit
```
  f62bd3b4
- Y
  Support bfloat16 for adamw and adam optimizer. Fit the lr for pure bf16... · e5ed5257
  由 Yuang Liu 提交于 11月 17, 2022
```
Support bfloat16 for adamw and adam optimizer. Fit the lr for pure bf16 training with tensor fusion. (#48041)

* add bfloat16 for adamw

* set lr not to bfloat16 for pure bf16 training

* update the logic

* update the adamw optimizer

* support bfloat for adam
```
  e5ed5257
- S
  Add vectorized bfloat16 atomicAdd (#48056) · ccbd03d5
  由 sneaxiy 提交于 11月 17, 2022
```
* add vectorized bfloat16 atomicAdd

* fix compile error

* fix compile error again

* fix V100 compile error

* fix V100 compile again
```
  ccbd03d5
- Z
  
  generate static graph code for some op (#48036) · 7cc0d171
  由 zyfncg 提交于 11月 17, 2022
  
  7cc0d171
16 11月, 2022 14 次提交
- H
  
  rm "paddle/fluid/framework/gpu_utils.h" in phi (#48020) · 29a0987a
  由 huangjiyi 提交于 11月 16, 2022
  
  29a0987a
- Q
  [NPU] update npu prop, test=develop (#47859) · ad8847aa
  由 Qi Li 提交于 11月 16, 2022
```
* [NPU] update npu prop, test=develop

* remove ddim.h

* remove diff

* update storage prop, test=develop
```
  ad8847aa
- X
  [Paddle Inference] Add fill_any_like trt converter. (#47974) · d6be9000
  由 xiaoxiaohehe001 提交于 11月 16, 2022
```
* add_fill_any_like

* add_fill_any_like
```
  d6be9000
- W
  elementwise_floordiv (#47944) · b4b78060
  由 wenbin 提交于 11月 16, 2022
```
* elementwise_op

* add teller

* modify ut

* comments

* modify ut

* return

* modify
```
  b4b78060
- Z
  
  trt memory set change from setMaxWorkspaceSize to setMemoryPoolLimit since trt 8.3+ (#47795) · 9cf3aa61
  由 Zhang Jun 提交于 11月 16, 2022
  
  9cf3aa61
- Z
  
  [inference][trt] update trt hardswish plugin to layer (#47745) · 6c54e0e8
  由 Zhang Jun 提交于 11月 16, 2022
  
  6c54e0e8
- H
  [Opt depthwise_conv2d] Simplify depthwise_conv2d use_cudnn attribute (#48010) · 7c304580
  由 HongyuJia 提交于 11月 16, 2022
```
* simplify depthwise_conv2d phi kernel selection

* fix depthwise_conv2d
```
  7c304580
- P
  Add bf16 data type support to oneDNN bilinear_interp kernel (#46770) · 8e6315e4
  由 Piotr Paturej 提交于 11月 16, 2022
```
* Enable bf16 in oneDNN bilinear_interp kernel

* Fix bilinear_interp_v2 not enabled in models

* Remove unnecessary checks
```
  8e6315e4
- Y
  Fix paddle rec, kim, dsin models' bugs (#47792) · e23dfed9
  由 ykkk2333 提交于 11月 16, 2022
```
* add stat tool

* add roll and roll_grad kernels and strided_slice and strided_slice_grad kernels, test=kunlun

* embedding and embedding_grad add int32 input, test=kunlun
```
  e23dfed9
- H
  remove avx check (#48003) · a762d68e
  由 hong 提交于 11月 16, 2022
```
* remove avx check

* fix bug;
```
  a762d68e
- L
  
  increase the level of some log (#47990) · 2f8901cb
  由 Leo Chen 提交于 11月 16, 2022
  
  2f8901cb
- W
  
  move "gpu_primitives.h" to phi (#48015) · 9adca1e7
  由 Wang Xin 提交于 11月 16, 2022
  
  9adca1e7
- W
  Update `ProcessGroupCustom` for `sync_op` compatibility (#47976) · e4ebf383
  由 Wen Sun 提交于 11月 16, 2022
```
* refactor: update pg custom

* fix: use new api in ut

* fix: typo

* revert: recover legacy apis

* fix: add GetDeviceContext
```
  e4ebf383
- C
  
  feat(ipu): add paddle inference support for model_runtime. (#47364) · 39c85064
  由 czr-gc 提交于 11月 16, 2022
  
  39c85064

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功