提交 · 0e51f3988bf92c3a13f3a1e54c0ade4d98c7edeb · PaddlePaddle / Paddle

31 1月, 2023 26 次提交
- C
  Integrate static code gen info (#49858) · 0e51f398
  由 Charles-hit 提交于 1月 31, 2023
```
* polish static grad op maker gen

* fix some bugs

* fix static code gen

* solve conflict

* modify composite grad maker name

* integrate phi and fluid info in static code gen

* rename some composite maker

* modify static code gen format
```
  0e51f398
- Z
  
  [inference][trt] add elementwise input data type check (#49675) · 5822e15c
  由 Zhang Jun 提交于 1月 31, 2023
  
  5822e15c
- P
  [Numpy] Add FP16 dtype for CastNumpy2Scalar (#50002) · 86a23818
  由 PuQing 提交于 1月 31, 2023
```
* add FP16 dtype for CastNumpy2Scalar

* fix throw message

* add test

* fix SyntaxWarning

* test skip for float16

* fix dtype mistakes
```
  86a23818
- R
  Add unified device management api (#48651) · 7aaaa1c6
  由 ronnywang 提交于 1月 31, 2023
```
* [CustomDevice] add custom device api

* update

* update

* test=document_fix

* update

* update

* add  examples
```
  7aaaa1c6
- R
  Fix 空指针 (Null pointer) of case15: paddle.broadcast_tensors (#49980) · 78ec942b
  由 RedContritio 提交于 1月 31, 2023
```
* fix incorrect output shape of broadcast

* add unittest
```
  78ec942b
- R
  
  fix send start msg (#50085) · 1048b166
  由 Roc 提交于 1月 31, 2023
  
  1048b166
- Z
  
  optimize 2D sync_batch_norm (#49663) · 9a4acfee
  由 zhangkaihuo 提交于 1月 31, 2023
  
  9a4acfee
- Z
  
  not use shm cache default (#50089) · 118aee6f
  由 zhangbo9674 提交于 1月 31, 2023
  
  118aee6f
- Bump Cutlass version to 2.11.0 (#50073) · c64296bf
  由 MarDino 提交于 1月 31, 2023
  
  c64296bf
- 张
  fix div 0 error in floormod (#49997) · 26bdea0f
  由张春乔提交于 1月 31, 2023
```
* fix mod 0 error

* fix div 0 error in floormod
```
  26bdea0f
- Y
  
  [Paddle Inference] change the default values of some gflags (#50074) · a1f28a48
  由 Yuanle Liu 提交于 1月 31, 2023
  
  a1f28a48
- 2
  
  support fp16 squaredl2norm (#48315) · ce4637c1
  由 201716010711 提交于 1月 30, 2023
  
  ce4637c1
- X
  support 0d tensor for interpolate (#49929) · 2e156ac8
  由 xiaoting 提交于 1月 31, 2023
```
* support 0d tensor for interpolate

* support 0d tensor for interpolate

* add xpu unittest for interp

* update unittest for interpolate

* fix coverage

* fix code style

* fix for coverage

* fix coverage
```
  2e156ac8
- T
  support inplaced variable in cinn_launch (#49912) · 754ab705
  由 TeFeng Chen 提交于 1月 31, 2023
```
* support inplaced variable in cinn_launch

* fix error hint when compiling

* fix inplaced output variable of the subgraph

* skip CinnCompiler check

* using existed definition

* fix namespace reference error

* modify error message

* update cinn tage

* fix namespace

* skip enforce check

* fix unittest attribute throw
```
  754ab705
- P
  
  change no_event GC to fast GC for xpu (#49871) · eba7b584
  由 pangyoki 提交于 1月 31, 2023
  
  eba7b584
- H
  [Decouple phi] Decouple custom_op in fluid and phi (#49866) · 48b3e869
  由 HongyuJia 提交于 1月 31, 2023
```
* decouple phi custom_op

* decouple phi custom_op, remove codes

* delete custom symbol of inference
```
  48b3e869
- 张
  
  fix div 0 error in conv1_transpose (#50000) · 1755a154
  由张春乔提交于 1月 31, 2023
  
  1755a154
- R
  Fix 堆栈溢出 (stack overflow) of case10: paddle.unique (#49981) · dbfdefa7
  由 RedContritio 提交于 1月 31, 2023
```
* add axis check in UniqueRawInferMeta

* add unittest for negative axis

* simplify check for unique
```
  dbfdefa7
- R
  Fix 空指针 (Null pointer) of case 14 paddle.atan2 (#49973) · 82edc65b
  由 RedContritio 提交于 1月 31, 2023
```
* add elements count check in atan2

* add unittest and pre-check in inferMeta

* add dimension check
```
  82edc65b
- 张
  Fix the div 0 error of matrix_power (#49942) · fb74147c
  由张春乔提交于 1月 31, 2023
```
* add zero size check in matrix_power_kernel_impl.h

* add zero size check in matrix_power_kernel_impl.h

* add zero size check in unittest

* bug_fix

* bug_fix

* bug_fix

* bug_fix

* bug_fix

* bug fix

* bug_fix

* bug_fix

* add static check

* delete the dy codes
```
  fb74147c
- R
  Fix 堆栈溢出 (stack overflow) of case9: paddle.repeat_interleave (#49982) · 66682be0
  由 RedContritio 提交于 1月 31, 2023
```
* support negative index in repeat_interleave

* add unittest
```
  66682be0
- 张
  
  fix the div 0 error of pixel_shuffle (#49996) · baf96a12
  由张春乔提交于 1月 31, 2023
  
  baf96a12
- R
  
  add dims check for nms_kernel (#49993) · 4976153d
  由 RedContritio 提交于 1月 31, 2023
  
  4976153d
- Y
  Unify the gpu implementation of stack and unstack to reuse the optimization. (#49748) · 3586e856
  由 Yiqun Liu 提交于 1月 31, 2023
```
* Unify the gpu implementation of stack and unstack to reuse the optimization.

* Optimize the cuda implementation of unstack.

* Use GpuMemcpyAsync instead of memory::Copy.

* Fix error of calculating the index.

* Use FastDivMod to further imporve the performance of unstack.
```
  3586e856
- L
  
  add multi fetch (#50070) · a8078bbd
  由 LiYuRio 提交于 1月 31, 2023
  
  a8078bbd
- 姜
  rm flags retain grad in pybind (#49888) · 9c3a35b9
  由姜永久提交于 1月 31, 2023
```
* rm flags_retain grad in pybind

* retain grads for xpu test

* set retain grad for xpu

* rm flag

* lint

---------
Co-authored-by: Nwanghuancoder <wanghuan29@baidu.com>
```
  9c3a35b9
30 1月, 2023 8 次提交

J

[CINN] fix build_cinn_pass collect inplace var bug (#50072) · ac84dce9
由 jiangcheng 提交于 1月 30, 2023

ac84dce9

Fix 空指针 (Null pointer) of case 2 paddle.linalg.lu_unpack (#49976) · 6f8ec229

由 RedContritio 提交于 1月 30, 2023

* add pivots type check and fix batchsize error

* add unittest for batchsize = 0

* fix nullptr in lu_unpack

fix batchsize error in LU_Unpack
add nullptr check in OneFunctor

* remove exception in device code

6f8ec229

[Divide by 0 Error] add pinv check (#49951) · f6e874bc

由 Ryan 提交于 1月 30, 2023

* add pinv check

* add unitest

* update unitest

* roll back

* fix not call stupid bug

* use context

f6e874bc

E
add phi tensor vector array api from fluid (#49885) · 094e3b8c
由 engineer1109 提交于 1月 30, 2023
```
replace all TensorFromVector & TensorToVector

AssignKernel async copy
```
094e3b8c

Support stream priority for standalone executor (#49939) · 172d1de6

由 Ruibiao Chen 提交于 1月 30, 2023

* Support stream priority for standalone executor

* Fix compile error

* Fix compile error

* Fix compile error

* Fix compile error

* Fix compile error

172d1de6

[Pglbox2.0] merge gpugraph to develop (#49946) · cb525d4e

由 zmxdream 提交于 1月 30, 2023

* add set slot_num for psgpuwraper (#177)

* add set slot_num_for_pull_feature for psgpuwarper

* Add get_epoch_finish python interface (#182)

* add get_epoch_finish interface

* add return

* delete return

* add unzip op (#183)

* fix miss key for error dataset (#186)

* fix miss key for error dataset

* fix miss key for error dataset
Co-authored-by: Nyangjunchao <yangjunchao@baidu.com>

* add excluded_train_pair and infer_node_type (#187)

* support return of degree (#188)

* fix task stuck in barrier (#189)
Co-authored-by: Nyangjunchao <yangjunchao@baidu.com>

* check node/feature format when loading (#190)

* check node&feature format when loading

* check node&feature format when loading (2£ (2)

* degrade log (#191)

* [PGLBOX]fix conflict

* [PGLBOX]fix conflict

* [PGLBOX]replace LodTensor with phi::DenseTensor

* [PGLBOX]fix gpu_primitives.h include path

* [PGLBOX]from platform::PADDLE_CUDA_NUM_THREADS to phi::PADDLE_CUDA_NUM_THREADS

* [PGLBOX]fix unzip example code

* [PGLBOX]fix unzip example code

* [PGLBOX]fix unzip example code

* [PGLBOX]fix unzip example code

* [PGLBOX]fix unzip ut

* [PGLBOX]fix unzip ut

* [PGLBOX]fix code style

* [PGLBOX]fix code style

* [PGLBOX]fix code style

* fix code style

* fix code style

* fix unzip ut

* fix unzip ut

* fix unzip ut

* fix unzip

* fix code stype

* add ut

* add c++ ut & fix train_mode_ set

* fix load into memory

* fix c++ ut

* fix c++ ut

* fix c++ ut

* fix c++ ut

* fix code style

* fix collective

* fix unzip_op.cc

* fix barrier

* fix code style

* fix barrier

* fix barrier

* fix code styple

* fix unzip

* add unzip.py

* add unzip.py

* fix unzip.py

---------
Co-authored-by: Nchao9527 <33347532+chao9527@users.noreply.github.com>
Co-authored-by: NSiming Dai <908660116@qq.com>
Co-authored-by: Nhuwei02 <53012141+huwei02@users.noreply.github.com>
Co-authored-by: Nyangjunchao <yangjunchao@baidu.com>

cb525d4e

G

depthwise_conv 映射成 conv的逻辑中添加下cudnn版本的判断 (#50058) · 320958eb
由 gem5 提交于 1月 30, 2023

320958eb
S
make FLAGS_gemm_use_half_precision_compute_type=false by default (#50050) · 964cd660
由 sneaxiy 提交于 1月 30, 2023
```
* make FLAGS_gemm_use_half_precision_compute_type=false defaultly

* fix comments
```
964cd660

29 1月, 2023 6 次提交
- J
  
  [CINN] BuildCinnPass collect inplace var from all cluster instead op (#50057) · 6d13992e
  由 jiangcheng 提交于 1月 29, 2023
  
  6d13992e
- Z
  
  refine code (#50053) · f8557cd9
  由 zhangbo9674 提交于 1月 29, 2023
  
  f8557cd9
- S
  
  update latest ps.proto (#50054) · 3da73f8f
  由 sneaxiy 提交于 1月 29, 2023
  
  3da73f8f
- S
  Add the missing ps.proto and remove ps_pb2.py (#50040) · ba67361b
  由 sneaxiy 提交于 1月 29, 2023
```
* add missing proto file

* fix windows ci

* fix ci compile error
```
  ba67361b
- R
  [CustomDevice] registering feed_dense_tensor, feed_sparse_coo_tensor,... · 50d92531
  由 ronnywang 提交于 1月 29, 2023
```
[CustomDevice] registering feed_dense_tensor, feed_sparse_coo_tensor, feed_strings kernels for custom device (#50042)

* [CustomDevice] registering feed_dense_tensor, feed_sparse_coo_tensor, feed_strings kernels for custom device

* update

* update

* update
```
  50d92531
- L
  [FleetExecutor] Remove max_slot_num and implement multi-scope fetch (#50041) · decbb588
  由 LiYuRio 提交于 1月 29, 2023
```
* remove max_slot_num

* fix test case
```
  decbb588

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功