提交 · fb7590d487431ba3b7b26bd3e0267c7194a127ff · Crayon鑫 / Paddle

25 4月, 2021 2 次提交

L
[NPU] refine lookup_table_v2_grad npu_kernel (#32497) · fb7590d4
由 Leo Chen 提交于 4月 25, 2021
```
* use ZerosLike instead of NPUMemsetAsync

* fix compile
```
fb7590d4

由 denglin-github 提交于 4月 25, 2021

* Add dlnne engine runtime

* Fix log

* Remove <const_cast> and remove unrelated modify with dlnne, +clang-format

* Fix CMakeList format error

* Add copyright message

* Fix dlnne CMakeList.txt

* Add some paddlepaddle_pass to support more networks

* Fix some format bug

feb2e476

24 4月, 2021 2 次提交
- Z
  
  clear CUDA compile environment on windows (#32498) · f8caa587
  由 Zhou Wei 提交于 4月 24, 2021
  
  f8caa587
- W
  
  refator paddle inference c api.test=develop (#32225) · 18d3e2ca
  由 winter-wang 提交于 4月 24, 2021
  
  18d3e2ca
23 4月, 2021 14 次提交
- C
  [CustomOp] Remove useless extension headers for old custom op (#32463) · 7d4998ac
  由 Chen Weihang 提交于 4月 23, 2021
```
* remove useless ext headers

* fix boost header compile failed
```
  7d4998ac
- L
  add the c_identity op (#32485) · 8fa8a37f
  由 lilong12 提交于 4月 23, 2021
```
* add c_identity op, test=develop
```
  8fa8a37f
- A
  Polish ParallelExectuor constructor into small functions (#32191) · faa8c703
  由 Aurelius84 提交于 4月 23, 2021
```
* Refine Constructor logic of ParallelExecutor

* refine function name

* refine code comment
```
  faa8c703
- L
  [NPU] refactor check_finite_and_scale npu kernel (#32407) · 39a59dcf
  由 Leo Chen 提交于 4月 23, 2021
```
* refactor_check_finite_and_scale_npu_kernel

* fix compile

* add alloc_float_status op

* add alloc_float_status op

* add FloatStatus for check_finite_and_unscale

* refine code

* remove unneccessary logic

* refine for fleet
```
  39a59dcf
- C
  
  ernie int8 support trt6 (#32424) · a01b5109
  由 ceci3 提交于 4月 23, 2021
  
  a01b5109
- W
  move semantic checks to op_teller (#32279) · 7c38114f
  由 wenbin 提交于 4月 23, 2021
```
* move semantic checks to op_teller

* more ops

* more ops

* revert block related change

* part1

* revert activation

* remove if

* remove const_cast

* reslove conflict

* remove const_cast

* delete useless var

* replace vlog(1) with vlog(3), replace assert with PADDLE_ENFORCE

* down to 19 files
```
  7c38114f
- Z
  fix Windows CI MP compile and environment install script and openblas CI (#32378) · 7a681f0b
  由 Zhou Wei 提交于 4月 23, 2021
```
* fix Windows CI MP compile and environment install script

* clear Windows CI environment

* clear Windows CI environment

* clear Windows CI environment
```
  7a681f0b
- B
  solve hccl communicate conflict (#32447) · 0e74eea2
  由 Baibaifan 提交于 4月 23, 2021
```
solve hccl communicate conflict (#32447)
```
  0e74eea2
- L
  add c_concat and c_split ops (#32486) · 2b108a04
  由 lilong12 提交于 4月 23, 2021
```
* add c_concat op
```
  2b108a04
- S
  
  add lstm support on xpu test=kunlun (#32436) · b6f8ccd2
  由 shanliang1992 提交于 4月 23, 2021
  
  b6f8ccd2
- W
  
  add WITH_STRIP=ON in paddle_build.sh, test=develop (#32450) · 51bcd97d
  由 wuhuanzhou 提交于 4月 23, 2021
  
  51bcd97d
- R
  
  [ROCM] add cuda kenrel for batch_norm_op (#32393) · 7879477f
  由 ronnywang 提交于 4月 23, 2021
  
  7879477f
- L
  
  [NPU] Fix bug that epsilon become 0 using power (#32469) · 49773f36
  由 Leo Chen 提交于 4月 23, 2021
  
  49773f36
- K
  Fix seven error message (#32397) · 203ac4f3
  由 Kqnonrime 提交于 4月 23, 2021
```
* fix two error message

* fix two error message

* fix error

* fix error

* fix error

* fix error

* fix some error message

* fix some error

* fix error

* fix some error

* fix some error

* fix some error

* fix one error

* fix some error

* fix seven error message

* fix error

* fix error

* fix error

* fix error
```
  203ac4f3
22 4月, 2021 7 次提交

W
support int32 and int64 kernel for clip operator (#32373) · c3328288
由 wuyefeilin 提交于 4月 22, 2021
```
support int32 and int64 kernel for clip operator 
```
c3328288
L

[NPU] remove ascend_parser for WITH_ASCEND_CL (#32451) · a1a527fb
由 Leo Chen 提交于 4月 22, 2021

a1a527fb
Z

Modify some contents for elementwise op impl (#32414) · 890d6bc0
由 Zhang Zheng 提交于 4月 22, 2021

890d6bc0
W

strip after compilation (#32145) · e727820d
由 wuhuanzhou 提交于 4月 22, 2021

e727820d

fix count problem (#32415) · 73d0b0e9

由 seemingwang 提交于 4月 22, 2021

* graph engine demo

* upload unsaved changes

* fix dependency error

* fix shard_num problem

* py client

* remove lock and graph-type

* add load direct graph

* add load direct graph

* add load direct graph

* batch random_sample

* batch_sample_k

* fix num_nodes size

* batch brpc

* batch brpc

* add test

* add test

* add load_nodes; change add_node function

* change sample return type to pair

* resolve conflict

* resolved conflict

* resolved conflict

* separate server and client

* merge pair type

* fix

* resolved conflict

* fixed segment fault; high-level VLOG for load edges and load nodes

* random_sample return 0

* rm useless loop

* test:load edge

* fix ret -1

* test: rm sample

* rm sample

* random_sample return future

* random_sample return int

* test fake node

* fixed here

* memory leak

* remove test code

* fix return problem

* add common_graph_table

* random sample node &test & change data-structure from linkedList to vector

* add common_graph_table

* sample with srand

* add node_types

* optimize nodes sample

* recover test

* random sample

* destruct weighted sampler

* GraphEdgeBlob

* WeightedGraphEdgeBlob to GraphEdgeBlob

* WeightedGraphEdgeBlob to GraphEdgeBlob

* pybind sample nodes api

* pull nodes with step

* fixed pull_graph_list bug; add test for pull_graph_list by step

* add graph table;name

* add graph table;name

* add pybind

* add pybind

* add FeatureNode

* add FeatureNode

* add FeatureNode Serialize

* add FeatureNode Serialize

* get_feat_node

* avoid local rpc

* fix get_node_feat

* fix get_node_feat

* remove log

* get_node_feat return  py:bytes

* merge develop with graph_engine

* fix threadpool.h head

* fix

* fix typo

* resolve conflict

* fix conflict

* recover lost content

* fix pybind of FeatureNode

* recover cmake

* recover tools

* resolve conflict

* resolve linking problem

* code style

* change test_server port

* fix code problems

* remove shard_num config

* remove redundent threads

* optimize start server

* remove logs

* fix code problems by reviewers' suggestions

* move graph files into a folder

* code style change

* remove graph operations from base table

* optimize get_feat function of graph engine

* fix long long count problem
Co-authored-by: NHuang Zhengjie <270018958@qq.com>
Co-authored-by: NWeiyue Su <weiyue.su@gmail.com>
Co-authored-by: Nsuweiyue <suweiyue@baidu.com>
Co-authored-by: Nluobin06 <luobin06@baidu.com>
Co-authored-by: Nliweibin02 <liweibin02@baidu.com>
Co-authored-by: Ntangwei12 <tangwei12@baidu.com>

73d0b0e9

support save/load binary format tensor. (#32211) · f4d9adc7

由 WeiXin 提交于 4月 22, 2021

* support save/load binary format tensor

* Fix error when create cudaplace

* Fix error when create cudaplace

* Fix error when create cudaplace

* get devive context from pool.

* move define of 'SerializeToStream' and 'DeserializeFromStream' to 'lod_tensor.cc' and 'selected_rows.cc'.

* improve coverage.

* improve coverage.

* polish API

* deal with conflict

* disable save/load large file in unnittest

* split unnittest.

f4d9adc7

T

Delete WITH_GRPC flag and Distributed old code (#32383) · e58c705b
由 tianshuo78520a 提交于 4月 22, 2021

e58c705b

21 4月, 2021 12 次提交

A

Add Bfloat16 support on Ampere GPU with CUDA 11 (#32132) · bf0ec9b8
由 AshburnLee 提交于 4月 21, 2021

bf0ec9b8

【NPU】Merge NPU ccl code (#32381) · c3158527

由 zhang wenhui 提交于 4月 21, 2021

* add allreduce and broadcast without test (#31024)

add allreduce and broadcast without test

* Refactor HCCLCommContext to be compatible with Paddle (#31359)

Refactor HCCLCommContext to be compatible with Paddle (#31359)

* [NPU] add npu kernel for communication op (#31437)

* add allreduce and broadcast without test

* add c_broadcast_test case

* build c_comm_init and c_create_group operators

* make the whole thing compile

* add broadcast and init op test case but run failed

* make unit test compile

* fix broadcast test bug and change into hcom for ccl

* change c_comm_init and c_create_group ops accordingly

* make tests compile

* transfer code to 27

* compiled successfully in 28, but run failed

* test broadcast in 28, but failed

* make hcom primitives work

* change hccl data type for base.h

* fix broadcast bug

* make attributes work

* fix group name bug

* add allreduce but test failed

* allreduce bug for qiuliang

* allreduce finished

* add allgather and reducescatter

* merge all op code

* add allgather test

* finish run all ccl op test exclude send/recv

* all all op and test exclude send/recv

* send_v2_npu.cc recv_v2_npiu.cc compiled

* fix ccl core dump bug and test allgather, reducescatter, broadcast op

* fix allreduce bug just for test

* hcom send&recv test pass, without hcom_destroy

* for qiuliang test

* Ascend Send&Recv Test Pass

* all op (ex send/recv) ok

* fix bug

* merge all ccl op

* style merge to PaddlePaddle

* merge style

* new merge style

* merge style 2

* insert an empty at the end

* disable ctest for hcom to pass ci
Co-authored-by: Nvoid-main <voidmain1313113@gmail.com>
Co-authored-by: Nf2hkop <f2huestc@outlook.com>

* Add auto-increasing tag id for Hcom OPs (#31702)

* add c_reduce_sum op (#31793)

add c_reduce_sum op

* update Ascendrc hccl to 20.3 (#32126)

update Ascendrc hccl to 20.3 (#32126)

* fix merge code

* change cmake.txt1

* [NPU] Support npu kernel for c sync stream op (#31386)

* sync stream npu op

* add with_ascend_acl

* update c++ unittest

* compile all failed

* try to pre commit

* after pre commit

* merge&compile&test hccl successfully!

* fix code style

* fix code style

* fix bugs about hccl

* fix some bugs

* fix code style

* fix style

* fix style

* fix

* fixed

* merge develop
Co-authored-by: Nlw921014 <liuwei921014@yeah.net>
Co-authored-by: NVoid Main <voidmain1313113@gmail.com>
Co-authored-by: Nf2hkop <f2huestc@outlook.com>
Co-authored-by: Nxiayanming <41795079@qq.com>

c3158527

C

Update the error info for quantizaion (#32273) · 3da2c7f3
由 cc 提交于 4月 21, 2021

3da2c7f3

optimize get-feat function of graph engine (#32261) · 2b68d20b

由 seemingwang 提交于 4月 21, 2021

* graph engine demo

* upload unsaved changes

* fix dependency error

* fix shard_num problem

* py client

* remove lock and graph-type

* add load direct graph

* add load direct graph

* add load direct graph

* batch random_sample

* batch_sample_k

* fix num_nodes size

* batch brpc

* batch brpc

* add test

* add test

* add load_nodes; change add_node function

* change sample return type to pair

* resolve conflict

* resolved conflict

* resolved conflict

* separate server and client

* merge pair type

* fix

* resolved conflict

* fixed segment fault; high-level VLOG for load edges and load nodes

* random_sample return 0

* rm useless loop

* test:load edge

* fix ret -1

* test: rm sample

* rm sample

* random_sample return future

* random_sample return int

* test fake node

* fixed here

* memory leak

* remove test code

* fix return problem

* add common_graph_table

* random sample node &test & change data-structure from linkedList to vector

* add common_graph_table

* sample with srand

* add node_types

* optimize nodes sample

* recover test

* random sample

* destruct weighted sampler

* GraphEdgeBlob

* WeightedGraphEdgeBlob to GraphEdgeBlob

* WeightedGraphEdgeBlob to GraphEdgeBlob

* pybind sample nodes api

* pull nodes with step

* fixed pull_graph_list bug; add test for pull_graph_list by step

* add graph table;name

* add graph table;name

* add pybind

* add pybind

* add FeatureNode

* add FeatureNode

* add FeatureNode Serialize

* add FeatureNode Serialize

* get_feat_node

* avoid local rpc

* fix get_node_feat

* fix get_node_feat

* remove log

* get_node_feat return  py:bytes

* merge develop with graph_engine

* fix threadpool.h head

* fix

* fix typo

* resolve conflict

* fix conflict

* recover lost content

* fix pybind of FeatureNode

* recover cmake

* recover tools

* resolve conflict

* resolve linking problem

* code style

* change test_server port

* fix code problems

* remove shard_num config

* remove redundent threads

* optimize start server

* remove logs

* fix code problems by reviewers' suggestions

* move graph files into a folder

* code style change

* remove graph operations from base table

* optimize get_feat function of graph engine
Co-authored-by: NHuang Zhengjie <270018958@qq.com>
Co-authored-by: NWeiyue Su <weiyue.su@gmail.com>
Co-authored-by: Nsuweiyue <suweiyue@baidu.com>
Co-authored-by: Nluobin06 <luobin06@baidu.com>
Co-authored-by: Nliweibin02 <liweibin02@baidu.com>
Co-authored-by: Ntangwei12 <tangwei12@baidu.com>

2b68d20b

L
[NPU] register npu finalize on exit (#32390) · 8e4c1936
由 Leo Chen 提交于 4月 21, 2021
```
* [NPU] register finalize on exit

* fix
```
8e4c1936

remove thrust include files (#32395) · ab6f8745

由 wuhuanzhou 提交于 4月 21, 2021

* remove thrust includes, test=develop

* fix compilation error, test=develop

* fix compilation of truncated_gaussian_random_op, test=develop

ab6f8745

L

[Kunlun]add collective ops for multi XPU cards training and add Kunlun multi XPU cards CI (#32302) · 2194ad15
由 liuyuhui 提交于 4月 21, 2021

2194ad15

石

flush denormal in the tracer op, test=develop (#32350) · 9ff85561

由石晓伟提交于 4月 21, 2021

* flush denormal in the tracer op, test=develop

* add cmake dependencies, test=develop

* add a macro, test=develop

* fix the windows case, test=develop

9ff85561

J

Added bilinear and nearest interp v2 oneDNN FP32 kernels (#32312) · 5d19f8d8
由 jakpiase 提交于 4月 21, 2021

5d19f8d8
I

Modify the exit code of mac CI approval error (#32389) · a2cbbe83
由 iducn 提交于 4月 21, 2021

a2cbbe83

add retry on gcda_clean.py (#32318) · 229f9308

由 YUNSHEN XIE 提交于 4月 21, 2021

* add retry on gcda_clean.py

* add exit code for paddle_coverage.sh

* fix format error

* fix format error

229f9308

J

Added oneDNN reduce_op GRAD kernel (#32280) · ead83422
由 jakpiase 提交于 4月 21, 2021

ead83422

20 4月, 2021 3 次提交

[Optimize]SparseKV speedup and memory save (#32048) · 5e7e7c9f

由 tangwei12 提交于 4月 20, 2021


Change-Id: Ie35a09772e46f7d90cb68ca82c1d18b9201d1abe

* large scale kv store optimize

Change-Id: I582cc661afdaa20749ec7493eae1b88c32b967f7

* replace std::unorded_map with roundrobin map

Change-Id: I48ee0efef38853876c92d982cdfcac6603c52c88

* remove license

* fix cpp lint

Change-Id: Ia21fafa65adc09bb9094f7dbc987e31d5af2686e

5e7e7c9f

W

move REGISTER_OP_CUDA_KERNEL into cpp with eigen, test=develop (#32114) · f6f59e50
由 wuhuanzhou 提交于 4月 20, 2021

f6f59e50
T
[heterps] optimize build task (#32358) · c09d6453
由 Thunderbrook 提交于 4月 20, 2021
```
* build task cost

* return pool
```
c09d6453

Crayon鑫 / Paddle 与 Fork 源项目一致

Crayon鑫 / Paddle
与 Fork 源项目一致