提交 · 8e5ed04d635b11e6234ec0dfa96c8e8133461574 · PaddlePaddle / Paddle

17 1月, 2023 17 次提交

support CUDA Graph for new executor (#49708) · 8e5ed04d

由 pangyoki 提交于 1月 17, 2023

* new exe supports CUDA Graph

* fix

* fix

* fix

* fix FLAGS_use_stream_safe_cuda_allocator in unittest

* insert output of coalesce_tensor op to skip_gc_var

* fix

8e5ed04d

P

add https://twitter.com/PaddlePaddle_ to README.md (#48227) · 76302bdc
由 PPGitub 提交于 1月 17, 2023

76302bdc

Prim api gen (#49654) · 813e27c9

由 xiaoguoguo626807 提交于 1月 17, 2023

* proto type of composite grad in paddle

* proto type of composite grad in paddle

* refactor composite api with phi

* fix compile error

* support static graph code-gen for squeeze op

* generate static graph code of unsqueeze

* refine op name

* fix compile error

* add extra output in op_compat

* remove debug log

* fix clang compile error

* support prim switch flag

* support prim switch flag

* fix dygraph error

* merge develop

* add code_gen

* add necessary files without codegen

* fix code_gen bug

* add deps

* modify igmnore

* add ignore

* delete std cout

* add composite logic for backward.py

* add tanh first order grad composite

* support enable_prim flag for static graph

* throw expection when both GrapOpMaker and GradCompOpMaker not been registered

* reorganize the directory of prim api tests

* fix windows error

* add eager_utils

* add eager_utils

* modify code gen

* add composite parse

* add unittest for get_grad_op_desc

* code optimize

* fix static test on windows

* support generate static graph code for imag and real op

* fix windows compile error in test_static_prim

* merge develop

* disable test eager in inference

* prim code gen

* disable eager compile in inference

* origin_yaml codegen success

* rm other file

* rm gitignore file

* code_style

* add eager test

* code_style

* clear #

* merge develop

* clear #

* remove useless files

* modify static test

* support bool flag from singlton

* merge develop

* recover git ignore

* fix conflict

* clear prim_gen

* recover git ignore for generated op

* parse_yaml success

* fix test compile error

* remove some tests

* add python test

* code_style

* revert parse_utils+ clear prim_gen

* fix some name issue

* add composite code gen

* modify backward yaml

* fix static composite grad maker code gen

* remove addtional files

* add some static funcs unit test

* fix some bugs

* fix composite grad maker register code gen

* optimize some functions

* modify gen cmake

* add more api gen

* add header

* modify static

* add static expand unsqueeze

* comments

* modify compopmaker

* revert

* modify gen name
Co-authored-by: NJiabinYang <360788950@qq.com>
Co-authored-by: Nzyfncg <zhangyunfei07@baidu.com>
Co-authored-by: Ncxxly <chenxx_id@163.com>
Co-authored-by: Ncharles-hit <wanghao107@baidu.com>

813e27c9

[PHI]Change feed_op to phi kernel (#49116) · f7f1dc03

由 YuanRisheng 提交于 1月 17, 2023

* change feed_op to phi kernel

* fix ci bugs

* fix build bugs

* fix ci bugs

* fix compile bugs

* fix ci bugs

* perfect code

* perfect comment code

* fix install bugs

* modify code according comment

* remove visitor in feed_op

* modify according comment

* perfect code according comment

* add infershape

* fix py3 bugs

* fix getexpected kernel type

* fix getexpected kernel type

* fix ci bugs

* add registry for custom device

* fix py3 bugs

* fix floating point error

* fix py3 test bugs

f7f1dc03

Merge ops composite into to_static (#49836) · b2a10916

由 cyber-pioneer 提交于 1月 17, 2023

* support @to_static+to_prime+cinn

* fix code logic

* debug4

* debug5

* debug6

* debug7

* debug 8

* debug 9

* debug10

* debug11

* debug11

* debug 12
Co-authored-by: NAurelius84 <zhangliujie@baidu.com>

b2a10916

W
Fix translated layer fine-tune error (#49870) · 412573f0
由 WangZhen 提交于 1月 17, 2023
```
* Fix translated layer fine-tune
```
412573f0
D

fix ps ut error;test=develop (#49867) · 56cacae9
由 danleifeng 提交于 1月 17, 2023

56cacae9
J

add test for composite with dy2st (#49873) · b927ce81
由 Jiabin Yang 提交于 1月 17, 2023

b927ce81

Support 0d Tensor in ConditionalBlockOp (#49842) · 791637cf

由 Huihuang Zheng 提交于 1月 17, 2023

Support 0d Tensor in ConditionalBlockOp

1. Add dygraph 0d tensor support for ConditionalBlockOp
2. Set scalar loss shape when `append_backward`

791637cf

姜

rm flag retain grad (#49835) · 73f97de0

由姜永久提交于 1月 17, 2023

* rm retain grad

* fix zero_dim

* fix zero_dim for xpu

* reset zero dim for xpu

* reset xpu

* reset custom_relu

* Reset flip

* fix zero dim

73f97de0

Fix build ci (#49879) · 60ee518a

由 risemeup1 提交于 1月 17, 2023

* fix build ci bug

* fix build ci bug,test=test=document_fix

* fix build ci bug,test=document_fix

60ee518a

Z

Fix the paddle/staitc/amp/__init__.py (#49791) · fcc90531
由 zhangkaihuo 提交于 1月 17, 2023

fcc90531
disable scatter zero_dim test (#49853) · 86fa1715
由 zhouweiwei2014 提交于 1月 17, 2023

86fa1715
H

SetDevice when parse TensorBase (#49860) · 4c576870
由 HongyuJia 提交于 1月 17, 2023

4c576870
W
[Dy2St]Support call backward() without params in dy2st (#49812) · 2f24b2d8
由 WangZhen 提交于 1月 17, 2023
```
* Support call backward() without params in dy2st
```
2f24b2d8
L

Modified compute and amplifier interceptor (#42044) · 989e39a5
由 LiYuRio 提交于 1月 17, 2023

989e39a5

【Prim】Add multiply,expand,div vjp rules (#49831) · 39c6765a

由 Xiaoxu Chen 提交于 1月 17, 2023

* support elementwise base func

* fix compiling error and add test

* support vjp for div using comp

* remove additional change

* fix dy2st error with magic num

* fix dy magic num

* another magic

* another magic

* another magic

* add skip rename strategy

* support add vjp

* support add with new axis cal

* support sub vjp

* [prim] add multiply vjp rules

* [prim] add multiply vjp rules

* [prim] fix no infershape with composite in _append_backward_ops

* [prim] add expand vjp rule

* [prim] add exp vjp rule

* uncomment infer shape for reshape/sum static prim api

* [prim] fix tanh nullptr error

* remove some print message

* fix magic number in run_program relative tests @JiaBinYang

* [prim] add expand,multiply,exp vjp rules

* fix only support single direction reduce error

* infer reduce dims using out dims
Co-authored-by: NJiabinYang <360788950@qq.com>

39c6765a

16 1月, 2023 22 次提交
- Support the 'data_transform' for generating static graph ops (#49772) · 28864137
  由 HappyHeavyRain 提交于 1月 16, 2023
```
* support the 'data_transform' for generating static graph ops

* reset 'pow' code

* change the 'GetKernelTypeForVar'
```
  28864137
- Z
  CUDA12.0 integration (#49539) · 1885d55a
  由 zlsh80826 提交于 1月 16, 2023
```
* Update warpctc for cuda-12

* Deprecate cudaProfilerInitialize for CUDA > 11

* Deprecate CUSPARSE_MV_ALG_DEFAULT for CUDA_VERSION >= 11040

* Add the missing thrust header
```
  1885d55a
- Z
  [inference] Use output var name to mark the NVTX flag (#49825) · ea2e2495
  由 Zhang Jun 提交于 1月 16, 2023
```
* add outvar name for nvtx mark

* nly network created with kEXPLICIT_BATCH can setsetMaxBatchSize
```
  ea2e2495
- W
  
  [PHI] channel_shuffle add yaml (#49808) · 56dbe426
  由 Weilong Wu 提交于 1月 16, 2023
  
  56dbe426
- W
  
  add add_n for the 0d tensor (#49854) · 65b0181e
  由 wawltor 提交于 1月 16, 2023
  
  65b0181e
- R
  
  optimize build_type (#49826) · 8fdb9087
  由 risemeup1 提交于 1月 16, 2023
  
  8fdb9087
- A
  [CINN]Switch cinn GIT_TAG from v0.2 into develop (#49775) · c8187ac7
  由 Aurelius84 提交于 1月 16, 2023
```
* [CINN]Switch cinn GIT_TAG from v0.2 into develop

* fix branch name

* specify commit

* disable unittest

* disable unittest
```
  c8187ac7
- Y
  [Paddle-TRT] support nhwc (#49633) · e43f7102
  由 Yuanle Liu 提交于 1月 16, 2023
```
* add trt_support_nhwc_pass
```
  e43f7102
- W
  
  [Fluid clean]clean distributed fluid API (#49795) · 7de9420a
  由 wangxiaoning 提交于 1月 16, 2023
  
  7de9420a
- W
  [fix code style]fix cpplint code style (#49742) · a3f58b70
  由 wangxiaoning 提交于 1月 16, 2023
```
* fix ctr_double_accessor.h

* fix graph_brpc_client.h non-const reference to pointer

* fix common_table.h

* fix graph_py_service.cc, server.cc, server.h
```
  a3f58b70
- G
  Fix paddle save for multi-processing (#49657) · 504db4f5
  由 Ghost Screaming 提交于 1月 16, 2023
```
* Fix bug of reduce_sum op. When input.numel() > INT32_MAX, its result
is wrong.

* Remove climits.

* Fix bug of paddle.save. It may cause bug for saving sharded optimizer
state_dict() in parallel.
```
  504db4f5
- W
  
  [fix code style]fix cpplint code style (#49811) · b7d44eb9
  由 wangxiaoning 提交于 1月 16, 2023
  
  b7d44eb9
- J
  Revert "[static code gen]Add phi and fluid info in static code gen (#49763)" (#49848) · 0355bb90
  由 Jiabin Yang 提交于 1月 16, 2023
```
This reverts commit 4d5265b8.
```
  0355bb90
- Q
  
  add prod for kunlun (#49816) · bd03652f
  由 QingshuChen 提交于 1月 16, 2023
  
  bd03652f
- [win] add windows cuda arch bin set (#49819) · 41230dc0
  由 zhouweiwei2014 提交于 1月 16, 2023
  
  41230dc0
- W
  
  [Eager] polish bmm api (#49823) · d3d69d8c
  由 Weilong Wu 提交于 1月 16, 2023
  
  d3d69d8c
- L
  
  fix ninja compilation for cinn (#49838) · 02384bc6
  由 Leo Chen 提交于 1月 16, 2023
  
  02384bc6
- Y
  add gpu_cpu_map_matmul_to_mul_pass to kGpuLowerPrecisionPasses (#49753) · 07514139
  由 Yuanle Liu 提交于 1月 16, 2023
```
* add gpu_cpu_map_matmul_to_mul_pass to kGpuLowerPrecisionPasses

* disable fc_elementwise_layernorm_fuse_pass in mixed precision
```
  07514139
- C
  [static code gen]Add phi and fluid info in static code gen (#49763) · 4d5265b8
  由 Charles-hit 提交于 1月 16, 2023
```
* polish static grad op maker gen

* fix some bugs

* fix static code gen

* solve conflict

* modify composite grad maker name
```
  4d5265b8
- Z
  
  add sqrt_comp_grad composite rule (#49769) · 70378584
  由 zqw_1997 提交于 1月 16, 2023
  
  70378584
- X
  
  【prim】vjp for reduce sum (#49736) · 292f3f77
  由 xiaoguoguo626807 提交于 1月 16, 2023
  
  292f3f77
- Y
  [Auto Parallel] Clear some fluid APIs (#49793) · e70af91d
  由 Yulong Ao 提交于 1月 16, 2023
```
* [Auto Parallel] Rename methods of ProcessMesh

* [Auto Parallel] Impl the python process_mesh by the c++ one

* [Auto Parallel] Add some minor modifications

* [Auto Parallel] Rename some methods

* [Auto Parallel] Remove unnecessary codes

* [Auto Parallel] Add back some removed files

* [Auto Parallel] Fix bugs

* [Auto Parallel] Fix a bug

* Update process_mesh.cc

* [Auto Parallel] Merge dist attrs of Python into C++

* [Auto Parallel] Add back deleted importing

* [Auto Parallel] Add back removed unittest

* [Auto Parallel] Remove type qualifiers of return types

* [Auto Parallel] Fix some bugs

* [Auto Parallel] Fix a bug of the quant pass

* [Auto Parallel] Fix the code style

* [Auto Parallel] Clear some fluid APIs
```
  e70af91d
15 1月, 2023 1 次提交

support mp on xpu (#49815) · 6a56bce7

由 Roc 提交于 1月 15, 2023

1 update xccl lib
2 when using comm_ctx, the allocator should be set manually.

6a56bce7

PaddlePaddle / Paddle 1 年多 前同步成功

PaddlePaddle / Paddle
1 年多前同步成功