提交 · 1b585b2896c05f08a69c6513ba16fd6817739118 · 机器未来 / Paddle

28 2月, 2022 7 次提交

由 seemingwang 提交于 2月 28, 2022

* graph engine demo

* upload unsaved changes

* fix dependency error

* fix shard_num problem

* py client

* remove lock and graph-type

* add load direct graph

* add load direct graph

* add load direct graph

* batch random_sample

* batch_sample_k

* fix num_nodes size

* batch brpc

* batch brpc

* add test

* add test

* add load_nodes; change add_node function

* change sample return type to pair

* resolve conflict

* resolved conflict

* resolved conflict

* separate server and client

* merge pair type

* fix

* resolved conflict

* fixed segment fault; high-level VLOG for load edges and load nodes

* random_sample return 0

* rm useless loop

* test:load edge

* fix ret -1

* test: rm sample

* rm sample

* random_sample return future

* random_sample return int

* test fake node

* fixed here

* memory leak

* remove test code

* fix return problem

* add common_graph_table

* random sample node &test & change data-structure from linkedList to vector

* add common_graph_table

* sample with srand

* add node_types

* optimize nodes sample

* recover test

* random sample

* destruct weighted sampler

* GraphEdgeBlob

* WeightedGraphEdgeBlob to GraphEdgeBlob

* WeightedGraphEdgeBlob to GraphEdgeBlob

* pybind sample nodes api

* pull nodes with step

* fixed pull_graph_list bug; add test for pull_graph_list by step

* add graph table;name

* add graph table;name

* add pybind

* add pybind

* add FeatureNode

* add FeatureNode

* add FeatureNode Serialize

* add FeatureNode Serialize

* get_feat_node

* avoid local rpc

* fix get_node_feat

* fix get_node_feat

* remove log

* get_node_feat return  py:bytes

* merge develop with graph_engine

* fix threadpool.h head

* fix

* fix typo

* resolve conflict

* fix conflict

* recover lost content

* fix pybind of FeatureNode

* recover cmake

* recover tools

* resolve conflict

* resolve linking problem

* code style

* change test_server port

* fix code problems

* remove shard_num config

* remove redundent threads

* optimize start server

* remove logs

* fix code problems by reviewers' suggestions

* move graph files into a folder

* code style change

* remove graph operations from base table

* optimize get_feat function of graph engine

* fix long long count problem

* remove redandunt graph files

* remove unused shell

* recover dropout_op_pass.h

* fix potential stack overflow when request number is too large & node add & node clear & node remove

* when sample k is larger than neigbor num, return directly

* using random seed generator of paddle to speed up

* fix bug of random sample k

* fix code style

* fix code style

* add remove graph to fleet_py.cc

* fix blocking_queue problem

* fix style

* fix

* recover capacity check

* add remove graph node; add set_feature

* add remove graph node; add set_feature

* add remove graph node; add set_feature

* add remove graph node; add set_feature

* fix distributed op combining problems

* optimize

* remove logs

* fix MultiSlotDataGenerator error

* cache for graph engine

* fix type compare error

* more test&fix thread terminating problem

* remove header

* change time interval of shrink

* use cache when sample nodes

* remove unused function

* change unique_ptr to shared_ptr

* simplify cache template

* cache api on client

* fix

* reduce sample threads when cache is not used

* reduce cache memory

* cache optimization

* remove test function

* remove extra fetch function

* graph-engine data transfer optimization

* support graph_split load&query

* remove logs

* change shards to pointer vector

* use inference

* remove test code

* renorm op

* simplify renorm op

* recover local changes

* recover renorm op kernel

* fix init

* add blanklines in renorm doc

* fix import

* fix import

* add renorm to init.py

* merge

* move index_sample op

* Delete api.h

* Delete api.cc

* fix

* remove logs

* recover infer shape of grad

* recover changes

* change shape

* fix label

* fix

* fix

* fix

* fix

* fix

* fix

* fix

* fix

* fix

* fix
Co-authored-by: NHuang Zhengjie <270018958@qq.com>
Co-authored-by: NWeiyue Su <weiyue.su@gmail.com>
Co-authored-by: Nsuweiyue <suweiyue@baidu.com>
Co-authored-by: Nluobin06 <luobin06@baidu.com>
Co-authored-by: Nliweibin02 <liweibin02@baidu.com>
Co-authored-by: Ntangwei12 <tangwei12@baidu.com>

1b585b28

0
[Phi]Move size, erfinv, pixel_shuffle infershape to phi (#39949) · a0cb3203
由 0x45f 提交于 2月 28, 2022
```
* move size, erfinv, pixel_shuffle infershape to phi

* fix erfinv infermeta
```
a0cb3203

Grid_sampler optimization (#39751) · 2c66775b

由 Lijunhui 提交于 2月 28, 2022

* init grid_sampler with mode=bilinear

* solve error

* rm fill constant

* rm head

* change block size

* change block size

* optimize

* apply existing config

2c66775b

[Phi] move truncated_gaussian_random kernel (#39971) · 23aa7a36

由 furnace 提交于 2月 28, 2022

* [Phi] move truncated_gaussian_random, copy kernels

* [Phi] move truncated_gaussian_random, kernel register

* [Phi] move truncated_gaussian_random, delete useless codes

23aa7a36

Z
PR-CI-Py3 change cpu test (#39659) · 3cb93edf
由 zhangchunle 提交于 2月 28, 2022
```
* update;test=cpu-py3
```
3cb93edf

[Pten->Phi PR4] Rename pten in funcs to phi (#39961) · eb42dd52

由 Chen Weihang 提交于 2月 28, 2022

* rename pten_utils to phi_utils

* rename pten_utils target

* rename Pten to Phi

* replace pten with phi

* resolve conflict

eb42dd52

[KP] Unify .cu and .xpu files with .kps files (#39917) · 0ff72e5d

由 Liu-xiandong 提交于 2月 28, 2022

* [KP] Unify .cu and .xpu files with .kps files

* fix CI bug in GPU and modify the list

* fix conflict

* modify the date

0ff72e5d

26 2月, 2022 4 次提交
- Y
  
  revert reshape op infershape (#39946) · b33a3c23
  由 YuanRisheng 提交于 2月 26, 2022
  
  b33a3c23
- F
  Move GumbelSoftmax OP to phi (#39873) · 581b2c64
  由 From00 提交于 2月 26, 2022
```
* Move GumbelSoftmax OP to phi

* platform::errors -> phi::errors; GumbelSoftmaxGradInferMeta -> backend.h/cc

* Use axis util in kernel impl

* Remove namespace platform::errors

* Use GetCPUEngine in Device Context
```
  581b2c64
- F
  Move BilinearTensorProduct OP to phi (#39903) · de8f2748
  由 From00 提交于 2月 26, 2022
```
* Move BilinearTensorProduct OP to phi

* Set dtype for Infermeta
```
  de8f2748
- C
  
  fix mkldnn softmax erro (#39951) · ab872efe
  由 Chen Weihang 提交于 2月 26, 2022
  
  ab872efe
25 2月, 2022 16 次提交
- F
  
  [phi] update code for mkl based fft (#39889) · 687902fc
  由 Feiyu Chan 提交于 2月 25, 2022
  
  687902fc
- J
  
  added logsoftmax oneDNN kernel (#39793) · 584844ec
  由 jakpiase 提交于 2月 25, 2022
  
  584844ec
- S
  Add MultiTensorApply to calculate L2-Norm in DistributedFusedLamb optimizer (#39900) · d32a0102
  由 sneaxiy 提交于 2月 25, 2022
```
* add multi tensor apply l2 norm

* add multi_tensor_apply code

* make sizeof(TensorMeta) smalller

* move code to distributed_fused_lamb_op.cu

* remove useless FLAGS
```
  d32a0102
- 0
  move eye、size、erfinv、pixel_shuffle OP to phi (#39712) · 639675de
  由 0x45f 提交于 2月 25, 2022
```
* move eye OP to pten

* move size OP to pten

* merge develop

* fix merge

* move files

* move erfinv OP to phi

* remove comment

* move pixel_shuffle OP to phi

* remove comment

* fix PT_REGISTER

* fix NPU

* fix CR

* remove size_sig.cc for PR-CI-Coverage
```
  639675de
- Z
  
  Fix conflict caused by wrong namespace (#39930) · d8fc7211
  由 Zhang Zheng 提交于 2月 25, 2022
  
  d8fc7211
- A
  [phi]migrate increment addmm multinomial cholesky InferShapes to phi (#39913) · 87b903a3
  由 Aganlengzi 提交于 2月 25, 2022
```
* [phi]migrate increment addmm multinomial cholesky InferShapes to phi

* set_dtype and mod MultinomialFunctor
```
  87b903a3
- L
  
  move diag_v2 to phi (#39914) · 783c4aba
  由 Linjie Chen 提交于 2月 25, 2022
  
  783c4aba
- Z
  
  replace implementation with cuda kernel (#39795) · 64f1485a
  由 Zhang Ting 提交于 2月 25, 2022
  
  64f1485a
- Z
  Optimize perf of softmax_with_cross_entropy (#39553) · bbe5228c
  由 Zhang Zheng 提交于 2月 25, 2022
```
* Optimize perf of softmax_with_cross_entropy

* fix

* fix

* fix accuracy error
```
  bbe5228c
- Z
  [bf16] add bf16 kernel: elementwise_add elementwise_mul elementwise_sub (#39716) · 2fedd39b
  由 zhangbo9674 提交于 2月 25, 2022
```
* add ele_add

* add ele_mul

* add ele_sub

* sovle conflict

* fix npu

* refine ele_add

* add ele_mul unittest

* refine ele_sub

* refine ci

* refine unittest
```
  2fedd39b
- F
  [Phi] mv kernel (#39861) · 2553af4f
  由 furnace 提交于 2月 25, 2022
```
[Phi] mv kernel 
```
  2553af4f
- J
  
  add reduce_min and reduce_max (#39899) · 44da9b42
  由 joeqiao12 提交于 2月 25, 2022
  
  44da9b42
- C
  [Phi] Support cudnn kernel moving & move softmax kernels (#39547) · 8895379a
  由 Chen Weihang 提交于 2月 25, 2022
```
* support cudnn kernel moving

* polish cmake rules

* add unittest for coverage

* remove orig kernel

* remove softmax cudnn kernel

* fix softmax test failed

* fix npu func error

* resolve conflict

* rename gpu dnn kernels

* fix name rule error

* fix compile error

* update fp16 namespace
```
  8895379a
- F
  
  [MLU] add elementwise_mul mlu kernel (#39864) · 04d324b2
  由 fwenguang 提交于 2月 25, 2022
  
  04d324b2
- N
  
  Fix a bug in IndexKernel data overflow (#39891) · 0615815d
  由 niuliling123 提交于 2月 25, 2022
  
  0615815d
- W
  
  fill_constant_batch_size_like op support fp16 (#39907) · b8cf8ca7
  由 WangXi 提交于 2月 25, 2022
  
  b8cf8ca7
24 2月, 2022 13 次提交

C
[PTen->Phi PR3] Rename pten make target to phi (#39832) · f77019a0
由 Chen Weihang 提交于 2月 24, 2022
```
* rename pten to phi

* fix infrt compile failed

* resolve conflict
```
f77019a0
Z

[MLU]add mlu kernel for allreduce (#39788) · ce207c3a
由 zn 提交于 2月 24, 2022

ce207c3a
A
[phi]migrate increment addmm multinomial cholesky kernels to phi (#39858) · b695fd95
由 Aganlengzi 提交于 2月 24, 2022
```
* migrate increment addmm multinomial cholesky kernels to phi

* test pr39869

* test pr39869

* fix style and ci
```
b695fd95
L
[phi] move randint to phi (#39872) · 127440c3
由 Leo Chen 提交于 2月 24, 2022
```
* move randint to phi

* use host generator
```
127440c3

build a Paddle Graph from CINN compiled program for execution with PE (#39724) · 4d042a83

由 TeFeng Chen 提交于 2月 24, 2022

* build a Paddle Graph from CINN compiled program for execution with PE

* update names of some variables

* fix random fail in build_cinn_pass_test and update some comments

* fix compiler error by merging phi pr

4d042a83

Optimize nearest_interp backward (#39067) · df0b4434

由 Lijunhui 提交于 2月 24, 2022

* nearest_interp_bw init

* optimize kernel config

* optimize kernel config

* fix struct init

* optimize code

* rm duplicated struct

df0b4434

[Phi]Move cross OP to phi (#39829) · 6c358a7c

由 0x45f 提交于 2月 24, 2022

* move cross forward OP

* move cross grad op to phi

* move infershape

* refine infershape

* rename ctx

* set dtype and layout in InferMeta

* refine code

6c358a7c

L
[phi] move bce_loss to phi (#39868) · 6fc5d88a
由 Linjie Chen 提交于 2月 24, 2022
```
* move bce_loss to phi

* refine PADDLE_ENFORCE

* revert PADDLE_ENFORCE

* fix ci
```
6fc5d88a
【Phi】Migrate poisson op into phi (#39814) · bbe441fc
由 zhouweiwei2014 提交于 2月 24, 2022
```
* Migrate poisson op into phi

* fix CI

* fix comment
```
bbe441fc

Added nearest interp v2 BF16 FWD kernel (#39490) · 2ec943a7

由 jakpiase 提交于 2月 24, 2022

* added nearest interp v2 bf16

* disabled bilinear interp nhwc test

* added skipping UT for gpu

* added NHWC support

* removed unnecessary statements

* minor change

* CI fix

* added appropriate changes to interpolate_v1

* fix after review

* minor change

* minor change

* revert unwanted deletions

* CI fix

2ec943a7

H
Optimize where_op and abs_grad_op by the elementwise interface (#39609) · c9699556
由 huangxu96 提交于 2月 24, 2022
```
* Optimize the where_op by the elementwise_op funtion

* Modified where_op & abs_grad_op by elementwise interface
```
c9699556
N

Fix a bug in IndexKernel out-of-memory (#39867) · 2136bd42
由 niuliling123 提交于 2月 24, 2022

2136bd42
L
optimize performance of lookup_table_v2_op (#39856) · d6038c22
由 Li Min 提交于 2月 24, 2022
```
* optimize block config  and fp16 atomicAdd perf for lookup_table_v2_grad.
```
d6038c22

机器未来 / Paddle 与 Fork 源项目一致

机器未来 / Paddle
与 Fork 源项目一致