提交 · bba13e219f9586ab52ef26d630c65144b2dfcf85 · PaddlePaddle / Paddle

18 8月, 2022 1 次提交

change to async mode for xpu multi-card training in static graph mode, test=kunlun (#45024) · 41bdf41d

由 zhangxiaoci 提交于 8月 18, 2022

* change to async mode for xpu multi-card training in static graph mode

* minor bugfix

* irrelevant. move to another pr

* move change to other pr

* fix stream issue

* fix 'stream not meet with current context' error

* fix branch diverge, test=kunlun

41bdf41d

10 8月, 2022 2 次提交
- Z
  add macro control in enforce_xpu.h, test=kunlun (#45022) · 9e74211f
  由 zhangxiaoci 提交于 8月 10, 2022
```
* add macro control in enforce_xpu.h, test=kunlun

* minor bugfix

* minor bugfix
```
  9e74211f
- L
  [new-exec] set cuda device before run (#44985) · 68b06ba6
  由 Leo Chen 提交于 8月 10, 2022
```
* set cuda device before run

* add header file

* fix compile
```
  68b06ba6
05 8月, 2022 1 次提交
- Q
  
  [DCU] fix hipDeviceAttributeManagedMemory not support on DTK, test=develop (#44816) · 075d7219
  由 Qi Li 提交于 8月 05, 2022
  
  075d7219
01 8月, 2022 2 次提交
- [Sparse] optimize sparse attention (#44743) · 1149a378
  由 zhouweiwei2014 提交于 8月 01, 2022
  
  1149a378
- W
  infer context fix place error. (#44726) · 74e46a93
  由 Wilber 提交于 8月 01, 2022
```
* infer context fix place error.

* update

* update
```
  74e46a93
29 7月, 2022 1 次提交

move CUDAStream to phi (#44529) · da3743fd

由 Leo Chen 提交于 7月 29, 2022

* init

* move CUDAStream to phi

* fix compilation

* merge develop

* add stream_owned_ member

* split cuda_stream.h

* fix cpu compile

* fix constructor

* fix bug

* fix windows compile

* fix inference test_levit

* fix windows tests

da3743fd

26 7月, 2022 2 次提交
- R
  
  [CustomDevice] add blas_axpby api for gradient_accumulator (#44584) · 0d51fcf1
  由 ronnywang 提交于 7月 26, 2022
  
  0d51fcf1
- W
  inference multi stream support handle lazy init. (#44563) · 1892a441
  由 Wilber 提交于 7月 26, 2022
```
* multi stream support handle lazy init.

* support eigen lazy init

* update

* fix ci problem
```
  1892a441
22 7月, 2022 1 次提交
- Y
  
  Add code of occupancy computing on DCU and avoid threadID bug for DCU profiler (#44520) · 8037901b
  由 yuguo 提交于 7月 22, 2022
  
  8037901b
20 7月, 2022 1 次提交
- L
  
  add eigen3 dependency for phi_backends (#44479) · fbfdea51
  由 Leo Chen 提交于 7月 20, 2022
  
  fbfdea51
19 7月, 2022 1 次提交

compile phi/backends into one static library (#44373) · 1047cb17

由 Leo Chen 提交于 7月 19, 2022

* compile into one static library

* fix xpu compile

* fix xpu compile

* fix inference compile

* fix inference compile

* add custom test

* revert one file

1047cb17

18 7月, 2022 2 次提交
- [Sparse] Add sparse matmul kernel(coo*dense->dense) (#44346) · 3f70b1d3
  由 zhouweiwei2014 提交于 7月 18, 2022
  
  3f70b1d3
- R
  
  [CustomDevice] remove unused file (#44358) · fd6dcdfe
  由 ronnywang 提交于 7月 18, 2022
  
  fd6dcdfe
15 7月, 2022 1 次提交
- Z
  support KL2 multi-card training, *test=kunlun (#43889) · 270f25e9
  由 zhangxiaoci 提交于 7月 15, 2022
```
* update xccl lib
    * use separate streams for compute/comm on XPU
    * add broadcast op to xpu2_op_list
```
  270f25e9
14 7月, 2022 2 次提交

[Phi]Improve the mechanism for mkldnn kernel in PHI (#43941) · e9b4d0be

由 YuanRisheng 提交于 7月 14, 2022

* adapt mkldnn kernel in PHI

* fix ci compile bugs

* fix compile bugs

* fix compile bugs

* fix compile bugs

* fix compile bugs

* delete comment

* fix compile bugs in windows-inference

* delete code for converage

* modify code by review

* modify code by review

* add todo

* fix compile bugs

* fix compile bugs

* fix compile bugs

* fix unittest bugsx

e9b4d0be

R
[CustomDevice] add custom ccl 1/2 (#44294) · d88e77a7
由 ronnywang 提交于 7月 14, 2022
```
* [CustomDevice] add custom ccl api

* add ut
```
d88e77a7

13 7月, 2022 1 次提交
- R
  [CustomKernel] capi add eager mode support (#44164) · 033ef5e9
  由 ronnywang 提交于 7月 13, 2022
```
* [CustomKernel] add capi eager mode support

* add ut

* add capi test
```
  033ef5e9
12 7月, 2022 1 次提交
- C
  [PHI] Clean glog header in public header (#44216) · b0c9f24a
  由 Chen Weihang 提交于 7月 12, 2022
```
* clean glog header in public header

* move marco pos
```
  b0c9f24a
06 7月, 2022 1 次提交
- H
  
  minor fix VLOG for xpu. test=kunlun. (#44099) · 502062da
  由 houj04 提交于 7月 06, 2022
  
  502062da
05 7月, 2022 1 次提交
- R
  Dataloader add custom device support (#44013) · a0dc361c
  由 ronnywang 提交于 7月 05, 2022
```
* Dataloader add custom device support

* update test=document_fix
```
  a0dc361c
02 7月, 2022 1 次提交

unify cpu context (#43989) · 09096aeb

由 Leo Chen 提交于 7月 01, 2022

* unify cpu context

* fix init()

* delete test_device_context

* fix test_scalar

09096aeb

28 6月, 2022 1 次提交
- 【Sparse】add SparseTensor mv kernel(csr*dense_vec->dence_vec, coo*dense_vec->dense_vec) (#43668) · 5161a047
  由 zhouweiwei2014 提交于 6月 28, 2022
```
* [Sparse]add SparseTensor mv kernel(csr*dense_vec->dence_vec, coo*dense_vec->dense_vec)

* fix CI
```
  5161a047
24 6月, 2022 2 次提交
- [Sparse] support batch compute of SparseTensor matmul/masked_matmul/softmax (#43703) · eec4e034
  由 zhouweiwei2014 提交于 6月 24, 2022
  
  eec4e034
- X
  
  change svd_cpu_kernel from Eigen to Lapack, speed up the compile from 120s -> 20s (#43784) · bafd8dec
  由 xiongkun 提交于 6月 24, 2022
  
  bafd8dec
18 6月, 2022 1 次提交
- remove unuse cuSparse function (#43626) · 4a08c781
  由 zhouweiwei2014 提交于 6月 18, 2022
  
  4a08c781
16 6月, 2022 1 次提交

[CustomKernel] add custom kernel c api (#42986) · 6fe10181

由 ronnywang 提交于 6月 16, 2022

* [CustomKernel] add custom kernel c api

* update

* update

* fix unable to export capi
Co-authored-by: Nronny1996 <524019753@qq.com>

6fe10181

15 6月, 2022 3 次提交

add some kernels(csr*dense->csr, dense*dense->csr) of SparseTensor matmul (#42935) · 346efe96
由 zhouweiwei2014 提交于 6月 15, 2022
```
* add some kernel(csr*dense->csr, dense*dense->csr) of SparseTensor matmul

* fix CI

* fix CI

* fix comment

* fix comment
```
346efe96

Use int64_t in GetGpuLaunchConfig1D and ElementwiseKernel as index type to... · 15577630

由 Yiqun Liu 提交于 6月 15, 2022

Use int64_t in GetGpuLaunchConfig1D and ElementwiseKernel as index type to support large tensor. (#43506)

* Change some data type from int to int64_t in GetGpuLaunchConfig1D to support large tensor.

* Use int64_t in ElementwiseKernel as index type to support large tensor.

15577630

R
Refactor dynload/port.h (#43431) · 332fdd1e
由 Ruibiao Chen 提交于 6月 15, 2022
```
* Refactor port.h

* Remove some unnecessary code

* Fix CI errors
```
332fdd1e

13 6月, 2022 2 次提交
- R
  
  Fix cmakelint errors for some files (#43428) · edf69ae0
  由 Ruibiao Chen 提交于 6月 13, 2022
  
  edf69ae0
- Z
  sparse convertion kernel support secondary dispatch (#43345) · 5752643b
  由 zhangkaihuo 提交于 6月 13, 2022
```
* use GpuMemcpy and GpuMemset

* sparse convert kernel support double dispatch by indices dtype

* cudaMemcpyKind->gpuMemcpyKind
```
  5752643b
09 6月, 2022 1 次提交
- M
  
  [sparse inference] Supporting 2:4 sparse inference (#43179) · 20b38cfa
  由 minghaoBD 提交于 6月 09, 2022
  
  20b38cfa
08 6月, 2022 1 次提交
- X
  
  call_once (#43206) · cad139a7
  由 xiaoxiaohehe001 提交于 6月 08, 2022
  
  cad139a7
07 6月, 2022 1 次提交
- W
  
  [multi-stream] Fix split and concat problem. (#43039) · 8c3777df
  由 Wilber 提交于 6月 07, 2022
  
  8c3777df
05 6月, 2022 1 次提交
- S
  
  【code format check upgrade】 step2：clang-format (#42840) · a3730dc8
  由 Sing_chan 提交于 6月 05, 2022
  
  a3730dc8
04 6月, 2022 1 次提交
- S
  
  【code format check upgrade】 step2：cmake-format (#43057) · 92568edb
  由 Sing_chan 提交于 6月 04, 2022
  
  92568edb
19 5月, 2022 1 次提交
- C
  [CompileOpt] Refine enforce code and remove boost/variant include (#41093) · ca359fec
  由 Chen Weihang 提交于 5月 19, 2022
```
* refine enforce code

* refine enforce code

* fix compile failed

* fix infrt failed
```
  ca359fec
13 5月, 2022 1 次提交
- W
  
  add gpu resources. (#42723) · 1280f294
  由 Wilber 提交于 5月 13, 2022
  
  1280f294
05 5月, 2022 1 次提交

update xpu depends (#42365) · d90e24ac

由 QingshuChen 提交于 5月 05, 2022

* update xpu depends
*test=kunlun

* minor
*test=kunlun
Co-authored-by: Nroot <root@yq01-sys-hic-p40-0091.yq01.baidu.com>

d90e24ac

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功