提交 · 0c7153a4f5f7fac7993aa0e5dce70ebca1c2fbb8 · PaddlePaddle / Paddle

28 12月, 2021 18 次提交

Utilize StreamSafeCUDAAllocator to support fast GC in new executor (#37642) · 0c7153a4

由 From00 提交于 12月 28, 2021

* fix reshape move storage error

* remove needless set type

* alloc tensor by shared storage

* Utilize StreamSafeCUDAAllocator to support fast GC in new executor

* Fix compile error for Windows and ROCm

* Fix compile error for Windows

* Modify UT stream_safe_cuda_alloc_test

* Modify UT stream_safe_cuda_alloc_test

* Rewrite fast GC

* Rewrite fast GC

* Fix compile error for BOOST_GET_CONST

* Fix compile error for BOOST_GET_CONST

* Changes default stream for StreamSafeCUDAAllocator

* Fix a small CI error

* Remove some redundant code

* Fix conflict

* Fix compile error for ROCm

* Fix Windoes CI error

* Fix CI error

* Remove some unnecessary code

* Fix CI error

* Add UT for fast GC

* Fix CI error

* add device-agnostic stream class

* add stream.h

* fix ut

* fix cpu compile

* Use RWLock in GetAllocator

* Fix CI error
Co-authored-by: NChen Weihang <chenweihang@baidu.com>
Co-authored-by: Nzhiqiu <chenqiuliang@baidu.com>

0c7153a4

add matmul_to_mul matmul_v2_to_mul matmul_v2_to_matmul test case (#37645) · bed71992

由 heliqi 提交于 12月 28, 2021

* add matmul_to_mul matmul_v2_to_mul matmul_v2_to_matmul test case

* modify skip func to ignore_pass_case func

* rebuild CI

* rebuild CI

* add test_map_xx_pass timeout

* add test_map_xx_pass timeout

* merge from develop

* add timeout notest;test=coverage

* Cmakelist add timeout

* add timeout

* add attr of matmul_v2

* add trt skip

* delete trt config

* add skip,  mul diff on 3080

bed71992

T

refine amax/amin example(#38525) · 00a50af8
由 Tao Luo 提交于 12月 28, 2021

00a50af8

Support test basic of Var and Layer (#38426) · 1fb80a6a

由 Jiabin Yang 提交于 12月 28, 2021

* Rearranged Eager AutoCodeGen directory structure

* Removed USE_OP in Eager AutoCodeGen

* Enabled generation for Operators without Grad/Inputs/Outputs

* Resolved operators without input

* Fixed merge conflicts

* Enabled Eager AutoCodeGen for 10+ more operators

* Refactored Eager AutoCodeGen with more organized helper objects

* Enabled Eager AutoCodeGen for operators with multiple OpBases

* Adjusted Eager AutoCodeGen to Enable Passing Output Tensor as Input Argument

* Handled Dispensable Inputs/Outputs in Eager AutoCodeGen

* Adjusted function generation/call between Python-C API & Dygraph API

* Synchronized auto-generated Python-C API with Dygraph Forward Functions

* support more eager tensor api

* fix merge compile error

* fix compile error and fit develop code

* support pure CPU

* fix some logic error in eager_mode

* support _varbase_creator in eager mode

* Added safe_initialized interface to EagerTensor for use in processing dispensable inputs

* for eager mode

* refine

* support multiple constructor for eager tensor

* add place related code

* polish code

* specific randint with dtype of int64

* Support pure cpu test

* eager logic

* refine test in pure cpu

* eager logic

* eager logic

* eager logic, test=develop

* skip core.eager when in inference, test=develop

* refine, test=develop

* refine, test=develop

* call RetainGrad after run forward kernel, test=develop

* refine, test=develop

* support dygraph util, meta, guard test

* support inference test

* refine test and fix initializer failed

* support create varbase and fix retain grad error

* fix windows error

* support test code coverage

* support test code coverage

* support test code coverage
Co-authored-by: Njim19930609 <jim19930609@gmail.com>
Co-authored-by: NWang Huan <wanghuan29@baidu.com>

1fb80a6a

W

fix ci problem (#38474) · 2e4cb279
由 Wilber 提交于 12月 28, 2021

2e4cb279
Z
refactor matmul directory in pten (#38227) · 982bf444
由 zyfncg 提交于 12月 28, 2021
```
* refactor matmul directory in pten

* fix merge conflict
```
982bf444

Add API and op for take_along_axis (#38396) · 3310f519

由 huangxu96 提交于 12月 28, 2021

* add API and op for take_along_axis

* fix compile dependency problem and add example code and doc

* add unitest

* delete some code for CI coverage

* fix code style problem

* fix as review

3310f519

G

fix adamw epsilon in cuda kernel (#37746) · 6f1bb3d6
由 Guoxia Wang 提交于 12月 28, 2021

6f1bb3d6
T
Add Amax and Amin API (#38417) · 340dfb26
由 Tao Luo 提交于 12月 28, 2021
```
* add amax/amin

* support axis is list
```
340dfb26

[pten] remove in_type arg in cast kernel (#38486) · 0637b9a6

由 chentianyu03 提交于 12月 28, 2021

* remove intype arg in cast kernel

* modify conj config in api.yaml by dictionary order

* rm unused code in cast_kernel.cu

0637b9a6

add reduce_prod_xpu. fix reduce_mean_xpu bug. (#38481) · 78836bb7

由 houj04 提交于 12月 28, 2021

* add reduce_prod_xpu. fix reduce_mean_xpu bug.

* iadd reduce_prod_xpu. fix reduce_mean_xpu bug. test=kunlun

78836bb7

L
[new-exec] add completion_nofifier (#38447) · 404a4a6a
由 Leo Chen 提交于 12月 28, 2021
```
* add completion_nofifier

* fix bug

* unregist event waiter
```
404a4a6a

add mul_lstm_fuse_pass ut (#37795) · 1db61c3e

由 baoachun 提交于 12月 28, 2021

* add mul_lstm_fuse_pass ut

* update mul_lstm_fuse_pass ut

* update ut

* update ut

* update ut

* add CPU ut cmake setting

* update ut

1db61c3e

Z
add pass base unittest (#38504) · ee5f3641
由 zhaoyingli 提交于 12月 28, 2021
```
* add pass base unittest

* update gpt model
```
ee5f3641
S

fix compile dir conflict with include_dirs (#38479) · e42ed7d1
由 sneaxiy 提交于 12月 28, 2021

e42ed7d1
Z

Fixed issue with offset,test=allcases (#38506) · dc30ad1d
由 Zhanlue Yang 提交于 12月 28, 2021

dc30ad1d

Fix scatter_op fp16 perf problem. (#38499) · 33ce249f

由 Li Min 提交于 12月 28, 2021

* Fix scatter_op fp16 perf problem.

* Add scatter into black list.

* Add scatter into black list for dygraph.

33ce249f

L

Add constructor for fused dropout param to ease use. (#38475) · f9e8a775
由 Li Min 提交于 12月 28, 2021

f9e8a775

27 12月, 2021 19 次提交

W

[fleet_executor] Add task loop thread pool (#38420) · dba59db7
由 WangXi 提交于 12月 27, 2021

dba59db7
fix english doc of some API (#38468) · 5b6b88ab
由 zhouweiwei2014 提交于 12月 27, 2021

5b6b88ab
S

fix bugs in fp16 for dp (#38405) · 1ab5c511
由 ShenLiang 提交于 12月 27, 2021

1ab5c511
Y
[PTen]move reshape kernel according to new directory (#38432) · 49216134
由 YuanRisheng 提交于 12月 27, 2021
```
* move reshape

* fix compile bugs

* delete manipulation file

* fix compile bugs
```
49216134
P
fix accumulator bug when multiple inplace OPs are executed continuously (#38406) · 113c8b93
由 pangyoki 提交于 12月 27, 2021
```
* fix accumulator bug

* fix unittest
```
113c8b93
Z
Refine clip_by_global_norm (#38209) · 65f7fa0d
由 zhangbo9674 提交于 12月 27, 2021
```
* refine clip

* delete unused code

* refine logic for clip
```
65f7fa0d
S
[BugFix]Fix bug in pfp16 in DataParallel (#38378) · e8e47581
由 ShenLiang 提交于 12月 27, 2021
```
* fix bug in pfp16

* fix hip

* fix hip
```
e8e47581
B

update mkldnn matmul_transpose_reshape fuse pass ut (#38467) · 9cfdae91
由 baoachun 提交于 12月 27, 2021

9cfdae91

add matmulv2_transpose_reshape_pass ut (#37416) · f664a533

由 baoachun 提交于 12月 27, 2021

* update mkldnn matmul_v2_transpose_reshape_fuse_pass ut

* update mkldnn matmul_v2_transpose_reshape_fuse_pass ut

* update ut

* update ut

f664a533

fix renorm (#38459) · b0c7144a

由 seemingwang 提交于 12月 27, 2021

* graph engine demo

* upload unsaved changes

* fix dependency error

* fix shard_num problem

* py client

* remove lock and graph-type

* add load direct graph

* add load direct graph

* add load direct graph

* batch random_sample

* batch_sample_k

* fix num_nodes size

* batch brpc

* batch brpc

* add test

* add test

* add load_nodes; change add_node function

* change sample return type to pair

* resolve conflict

* resolved conflict

* resolved conflict

* separate server and client

* merge pair type

* fix

* resolved conflict

* fixed segment fault; high-level VLOG for load edges and load nodes

* random_sample return 0

* rm useless loop

* test:load edge

* fix ret -1

* test: rm sample

* rm sample

* random_sample return future

* random_sample return int

* test fake node

* fixed here

* memory leak

* remove test code

* fix return problem

* add common_graph_table

* random sample node &test & change data-structure from linkedList to vector

* add common_graph_table

* sample with srand

* add node_types

* optimize nodes sample

* recover test

* random sample

* destruct weighted sampler

* GraphEdgeBlob

* WeightedGraphEdgeBlob to GraphEdgeBlob

* WeightedGraphEdgeBlob to GraphEdgeBlob

* pybind sample nodes api

* pull nodes with step

* fixed pull_graph_list bug; add test for pull_graph_list by step

* add graph table;name

* add graph table;name

* add pybind

* add pybind

* add FeatureNode

* add FeatureNode

* add FeatureNode Serialize

* add FeatureNode Serialize

* get_feat_node

* avoid local rpc

* fix get_node_feat

* fix get_node_feat

* remove log

* get_node_feat return  py:bytes

* merge develop with graph_engine

* fix threadpool.h head

* fix

* fix typo

* resolve conflict

* fix conflict

* recover lost content

* fix pybind of FeatureNode

* recover cmake

* recover tools

* resolve conflict

* resolve linking problem

* code style

* change test_server port

* fix code problems

* remove shard_num config

* remove redundent threads

* optimize start server

* remove logs

* fix code problems by reviewers' suggestions

* move graph files into a folder

* code style change

* remove graph operations from base table

* optimize get_feat function of graph engine

* fix long long count problem

* remove redandunt graph files

* remove unused shell

* recover dropout_op_pass.h

* fix potential stack overflow when request number is too large & node add & node clear & node remove

* when sample k is larger than neigbor num, return directly

* using random seed generator of paddle to speed up

* fix bug of random sample k

* fix code style

* fix code style

* add remove graph to fleet_py.cc

* fix blocking_queue problem

* fix style

* fix

* recover capacity check

* add remove graph node; add set_feature

* add remove graph node; add set_feature

* add remove graph node; add set_feature

* add remove graph node; add set_feature

* fix distributed op combining problems

* optimize

* remove logs

* fix MultiSlotDataGenerator error

* cache for graph engine

* fix type compare error

* more test&fix thread terminating problem

* remove header

* change time interval of shrink

* use cache when sample nodes

* remove unused function

* change unique_ptr to shared_ptr

* simplify cache template

* cache api on client

* fix

* reduce sample threads when cache is not used

* reduce cache memory

* cache optimization

* remove test function

* remove extra fetch function

* graph-engine data transfer optimization

* support graph_split load&query

* remove logs

* change shards to pointer vector

* use inference

* remove test code

* renorm op

* simplify renorm op

* recover local changes

* recover renorm op kernel

* fix init

* add blanklines in renorm doc

* fix import

* fix import

* add renorm to init.py
Co-authored-by: NHuang Zhengjie <270018958@qq.com>
Co-authored-by: NWeiyue Su <weiyue.su@gmail.com>
Co-authored-by: Nsuweiyue <suweiyue@baidu.com>
Co-authored-by: Nluobin06 <luobin06@baidu.com>
Co-authored-by: Nliweibin02 <liweibin02@baidu.com>
Co-authored-by: Ntangwei12 <tangwei12@baidu.com>

b0c7144a

L
add device-agnostic stream class (#38391) · 6b5e33b4
由 Leo Chen 提交于 12月 27, 2021
```
* add device-agnostic stream class

* add stream.h

* fix ut

* fix cpu compile
```
6b5e33b4
S

refine float16 implementation (#38439) · 78375990
由 sneaxiy 提交于 12月 27, 2021

78375990
S

refine CUDA Graph (#38401) · 5f7e4a21
由 sneaxiy 提交于 12月 27, 2021

5f7e4a21

Support multi-outputs feature for broadcast ops (#38329) · 89d38f55

由 limingshu 提交于 12月 27, 2021

* No harm to KP

* Pass the compile stage

* change the WriteData function

* fix template bugs and pass ctest of current elementwise

* for passing partial template specialization of tempalte function in CI-ROCm

* To make 'WriteData' funtion flexible.

* a less harmful way to support multi-output

* a less harmful way to support multi-output

89d38f55

C

remove npu related impl (#38428) · f1d56b77
由 Chen Weihang 提交于 12月 26, 2021

f1d56b77
C
[PTen] Move cast kernel impl (#38382) · 1fb734d7
由 Chen Weihang 提交于 12月 26, 2021
```
* rename to api to copy_to

* revert needless change

* polish format
```
1fb734d7
B

add attr check for infer in batch_norm_act mkldnn fuse pass (#38443) · 04527ee3
由 baoachun 提交于 12月 27, 2021

04527ee3
G

gelu using normcdf for cudnn (#38450) · 37022482
由 Guoxia Wang 提交于 12月 27, 2021

37022482
Z
[AMP] Fix amp.decorate bug: parameters for non leaf layers cannot be decotated (#38402) · 5d902954
由 zhangbo9674 提交于 12月 27, 2021
```
* fix bug

* refine code

* refine code

* refine code
```
5d902954

26 12月, 2021 3 次提交
- C
  [PTen] Move copy kernel impl (#38421) · 73819658
  由 Chen Weihang 提交于 12月 26, 2021
```
* add register general kernel marco

* move copy kernel impl

* revert needless change

* polish details

* fix xpu compil faild

* fix xpu compile failed

* polish format
```
  73819658
- C
  
  auto parse kernel deps by include (#38438) · e5c7ca48
  由 Chen Weihang 提交于 12月 26, 2021
  
  e5c7ca48
- Z
  
  improve forward performace (#38279) · acef85b2
  由 Zhang Ting 提交于 12月 26, 2021
  
  acef85b2

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功