提交 · 1cd7e68b34487f06fb52acf5381fdd533425b9e0 · Crayon鑫 / Paddle

25 8月, 2022 5 次提交

optimize conv algo cache (#41891) · 1cd7e68b

由 hong 提交于 8月 25, 2022

* optimizer conv alog speed

* code polish

* remove useless code

* fix compile error

* fix cpu compile error

* not use cudnn alog t

* add search cache max number

* polish code

* fix cache test bug

* add groups data format to conv args

* fix cache test bug

* fix cudnn_deterministic bug

* fix test switch auto tune bug

* fix test swith autotune bug;

* fix conv cache bug

* fix cache test error

* fix cache test bug

* fix windows mac compile error

* fix workspace search error

* update cudnn cache

* fix cache test bug; test=develop

* fix autotune swith test error

* polish code

* oplish code

1cd7e68b

R

[triu_indices] add triu_indices_op (#45168) · a410c397
由 Rayman 提交于 8月 25, 2022

a410c397
S
Fix unique_kernel bugs (#45032) · ea1f4702
由 sprouteer 提交于 8月 25, 2022
```
* fix unique_kernel bugs

* fix unique kernel cu bugs
```
ea1f4702

Fix relu python call (#45082) · 839fac65

由 hong 提交于 8月 25, 2022

* add python final state

* fix bug

* fix bugs

* fix bug

* fix bug

* revert impl, final state mul not support selected rows

* fix softmax use cudnn error

* add softlable false unitest

* revert loss.py

839fac65

H

add temporal shift and grad *test=kunlun (#45300) · 63d9a175
由 haosicheng 提交于 8月 25, 2022

63d9a175

24 8月, 2022 8 次提交

make tensor_util contains no cuda code (#45256) · 78916a7a

由 Leo Chen 提交于 8月 24, 2022

* make tensor_util contains no cuda code

* refine isfinite

* revert ut

* move isfinite function to its op

* fix test

* fix compile

* std::isnan is not defined for int type on windows

* fix windows compile

* fix fp16

* fix rocm compile

* revert gradient node

78916a7a

P
fix arm compile problem in phi/CmakeLists.txt (#42325) · 0d05e646
由 pangyoki 提交于 8月 24, 2022
```
* fix arm compile problem

* fix format

* fix ci error
```
0d05e646
C

remove needless comments (#45376) · 1badb4c6
由 Chen Weihang 提交于 8月 24, 2022

1badb4c6

[phi] Transfer merged_momentum yaml to phi (#45359) · 09acc860

由 HongyuJia 提交于 8月 24, 2022

* add legacy_api.yaml

* set merged_momentum inplace only

* support inplace optional<vector<tensor>>

* add dygraph_mode api

* optimize TensorToConstDenseTensorPtr

09acc860

W

Adapt tensor axis for cumsum (#45372) · 7f49b9ba
由 WangZhen 提交于 8月 24, 2022

7f49b9ba

【Hackathon No.34】优化 poisson op (#45160) · 3c14b094

由 Rayman 提交于 8月 24, 2022

* 【Hackathon No.34】优化 poisson op

* [poisson] code style fix

* modify code style

* prevent from big number

* modify code style

* modify code style

* modify import

* modify import

* modify code style

3c14b094

W
[OpAttr]Adapt tensor minlength for bincount (#45342) · 12917c8c
由 WangZhen 提交于 8月 24, 2022
```
* Adapt minlength attr for bincount
```
12917c8c

[ Dy2Static | Controlflow ]While + Cond support for python container. (#45105) · f8f66ec5

由 xiongkun 提交于 8月 24, 2022

* while support for python container.
It is convenient to convert more dynamic graph codes into static graphs.

* cond support python container

f8f66ec5

23 8月, 2022 8 次提交
- N
  
  Delete the template parameter BLockSize in Kernel Primitive API (#45220) · 1a0cd447
  由 niuliling123 提交于 8月 23, 2022
  
  1a0cd447
- Z
  [Sparse]Use shorted function names (#45325) · 3a7b1810
  由 zhangkaihuo 提交于 8月 23, 2022
```
* rename the member function of SparseTensor

* use shorter function names
```
  3a7b1810
- L
  
  first commit (#45253) · b5d8bd2f
  由 limingshu 提交于 8月 23, 2022
  
  b5d8bd2f
- S
  
  [Geometric] Fix cuda configuration error for message_passing api (#45315) · 03ef0bdc
  由 Siming Dai 提交于 8月 23, 2022
  
  03ef0bdc
- T
  【PaddlePaddle Hackathon 3 No.33】为 Paddle 优化 erfinv op 在 GPU 上的计算性能 (#45057) · 0e384ade
  由 thunder95 提交于 8月 23, 2022
```
* erfinv

* fix some tiny issues
```
  0e384ade
- O
  modify something unimportant when I read source code (#45273) · 5edc96e6
  由 OccupyMars2025 提交于 8月 23, 2022
```
* Update scope.h

* typo

* Update dense_tensor.inl
```
  5edc96e6
- Y
  [Phi]Move distribute_fpn_proposals to PHI (#45212) · 8f8ed7de
  由 YuanRisheng 提交于 8月 23, 2022
```
* move distribute_fpn_proposals

* fix some code

* fix yaml bugs

* add set dtype

* move proposal_impl to funcs

* fix compile bugs
```
  8f8ed7de
- R
  [CustomDevice] add profiler apis (#45130) · da51baf2
  由 ronnywang 提交于 8月 23, 2022
```
* [CustomDevice] add profiler apis

* migrate CalculateEstOccupancy into cuda_tracer

* update

* add ut
```
  da51baf2
22 8月, 2022 4 次提交
- W
  [Eager] some python c api use final state (#45221) · d2ef888b
  由 wanghuancoder 提交于 8月 22, 2022
```
some python c api use final state
```
  d2ef888b
- Z
  
  rename the member function of SparseTensor (#45291) · 016b94c2
  由 zhangkaihuo 提交于 8月 22, 2022
  
  016b94c2
- S
  
  fix infershape in compile time (#45156) · ed57237e
  由 shangliang Xu 提交于 8月 22, 2022
  
  ed57237e
- R
  
  [CustomDevice] fix custom ccl (#45276) · 307ad60d
  由 ronnywang 提交于 8月 22, 2022
  
  307ad60d
19 8月, 2022 4 次提交
- P
  call final_state method in inplace APIs (#42968) · 7c1e7e46
  由 pangyoki 提交于 8月 19, 2022
```
* add forward inplace final state api

* fix bug

* fix reshape

* fix coverage

* add inplace info for erfinv, lerp, put_along_axis

* fix put_along_axis infer_meta

* fix format

* update yaml

* fix
```
  7c1e7e46
- H
  
  polish REGISTER_OPERATOR parameter of fill_any (#45263) · 1c4134f6
  由 HongyuJia 提交于 8月 19, 2022
  
  1c4134f6
- W
  Trt groupnorm dynamic plugin (#44911) · 1aa6adb1
  由 Wang Bojun 提交于 8月 19, 2022
```
* add group_norm dyanmic plugin
```
  1aa6adb1
- A
  
  [CustomDevice] support scalar (#45244) · dc331231
  由 Aganlengzi 提交于 8月 19, 2022
  
  dc331231
18 8月, 2022 6 次提交

[phi] Transfer fluid trilinear_interp_v2 to phi trilinear_interp (add yaml) (#45145) · 6150fade

由 HongyuJia 提交于 8月 18, 2022

* transfer trilinear op to phi, change name from trilinear_interp_v2 to trilinear_interp

* reserve linear_interp param

* change testcase scale if-branch

* testcase test_imperative_case

* fix trilinear testcase

* import paddle in test_trilinear_interp_v2

6150fade

A
[OpAttr]Squeeze axes support Tensor (#45189) · c93451f4
由 Aurelius84 提交于 8月 18, 2022
```
* [OpAttr]Squeeze axes support Tensor

* add support_tensor

* fix unittest

* fix coverage
```
c93451f4

change to async mode for xpu multi-card training in static graph mode, test=kunlun (#45024) · 41bdf41d

由 zhangxiaoci 提交于 8月 18, 2022

* change to async mode for xpu multi-card training in static graph mode

* minor bugfix

* irrelevant. move to another pr

* move change to other pr

* fix stream issue

* fix 'stream not meet with current context' error

* fix branch diverge, test=kunlun

41bdf41d

W

sync_batch_norm_backword_yaml (#45218) · 133f608f
由 wanghuancoder 提交于 8月 18, 2022

133f608f

[phi] Transfer fluid bilinear_interp_v2 to phi bilinear_interp (add yaml) (#45140) · 2c2137bb

由 HongyuJia 提交于 8月 18, 2022

* transfer bilinear op to phi, change bname from bilinear_interp_v2 to bilinear_interp

* reserve linear_interp param

* fix cross device import

2c2137bb

Z

support selected_rows kernel for multiply in dygraph (#45217) · bcbb7a97
由 zyfncg 提交于 8月 18, 2022

bcbb7a97

17 8月, 2022 4 次提交
- L
  Reuse addKernel to replace TensorAdd (#45161) · 0e3b49d4
  由 Leo Chen 提交于 8月 17, 2022
```
* use addKernel

* fix compile

* remove elementwiseAddto

* add return

* fix custom place
```
  0e3b49d4
- Y
  add instance norm op for xpu (#45097) · 216d25ac
  由 ykkk2333 提交于 8月 17, 2022
```
* xpu unittest grad compute supports more types, *test=kunlun

* add instance norm xpu, *test=kunlun
```
  216d25ac
- H
  [phi] Transfer fluid bicubic_interp_v2 to phi bicubic_interp (add yaml) (#45151) · f4da2d4d
  由 HongyuJia 提交于 8月 17, 2022
```
* transfer bicubic_interp op to phi, change name from bicubic_interp_v2 to bicubic_interp

* test final_state_bicubic_interp api

* testcase match imperative case
```
  f4da2d4d
- S
  Fix squared_l2_norm wrong stream bug (#45174) · 951010a2
  由 sneaxiy 提交于 8月 17, 2022
```
* fix squared_l2_norm bug

* update buffer.h
```
  951010a2
16 8月, 2022 1 次提交

[Phi] Move amp ops into phi (#45079) · b4f67757

由 Chen Weihang 提交于 8月 16, 2022

* move check finite and unscale kernel into phi

* move infershape into phi

* move update_loss_scaling kernel into phi

* remove original kernels

* move update loss scaling infershape into phi

* add header for xpu and npu

* solve coverage failed

* fix npu test failed

* remove mutable data in cu file

* fix new executor failed

* add valid check for meta tensor output

b4f67757

Crayon鑫 / Paddle 与 Fork 源项目一致

Crayon鑫 / Paddle
与 Fork 源项目一致