提交 · 266955a9e869ca325d7d74cd6c603744641aacd5 · BaiXuePrincess / Paddle

09 2月, 2022 6 次提交

C

move stream into pten (#39392) · 266955a9
由 Chen Weihang 提交于 2月 09, 2022

266955a9

update basic infrastructure (#39383) · b12e7a17

由 hong 提交于 2月 09, 2022

* update basic infrastructure; support string,  suport vecotr<int>, add tensor args type index; test=develop

* remove useless code; test=develop

* fix bug; test=develop

* polish code; test=develop

b12e7a17

Z

Modify the implementation of BlockYReduce to fit more scenes (#39170) · 8d87b3bc
由 Zhang Zheng 提交于 2月 09, 2022

8d87b3bc
N

Delete BASE_SIZE in elementwise_base.h (#39390) · b007a031
由 niuliling123 提交于 2月 09, 2022

b007a031

Add a Sparse Op: to_sparse_csr (#39333) · 76d527e1

由 zhangkaihuo 提交于 2月 09, 2022

* implement AllocateFrom

* dense_to_sparse_coo

* optimize unit testing; support rocm

* 1. delete fluid related header file
2. update the copyright

* fix hipMemcpy

* update dense_to_sparsecoo

* add namespace sparse

* sparse_csr_to_dense

* test to_sparse_coo: csr_to_coo

* fix writing error

* to_sparse_csr: dense_to_sparse_csr and sparse_coo_to_csr

* fix check shape

* fix unit test

* replace CUDADeviceContext by GPUContext

76d527e1

Move norm to pten (#39324) · ece200b3

由 hong 提交于 2月 09, 2022

* add norm cpu

* update code;

* norm bug fix

* move norm op to pten; test=develop

* move norm op to pten; test=develop

* add norm util; test=develop

* fix norm npu bug; test=develop

* fix norm kernel bug; test=develop

* move kernel args to pten; test=develop

* move kernel args to pten sig; test=develop

ece200b3

08 2月, 2022 7 次提交

Z
[bf16] add bf16 cuda kernel: concat and split (#39380) · de0bad2a
由 zhangbo9674 提交于 2月 08, 2022
```
* add concat & split

* add concat kernel

* add concat unittest

* add split unittest
```
de0bad2a
W
[PTEN] Update gpu_context. (#39359) · 24103cbb
由 Wilber 提交于 2月 08, 2022
```
* gpu_context..

* update

* update

* update
```
24103cbb
N
Replace clip, bce_loss, full and full_like with elementwise (#39197) · 424700ff
由 niuliling123 提交于 2月 08, 2022
```
* Replace clip, bce_loss, full and full_like with elementwise
```
424700ff

[PTen] Support SelectedRows in execution and remove scale OpKernel and InferShape (#39351) · 41eb2595

由 Chen Weihang 提交于 2月 08, 2022

* adapt selectedrows in execution

* impl selected rows branch

* support selectedrow in infershape utils

* fix device compile failed

* fix new exe test failed

* revert some changes

41eb2595

Support allocate CUDA managed memory (#39075) · 42910361

由 From00 提交于 2月 08, 2022

* Rough implementation for experiment

* Support allocate cuda managed memory

* Fix CI error

* Modify UT

* Check whether support memory oversubscription

* Fix ROCM Compile error

* Fix ROCM Compile error

* Fix UT cuda_managed_memory_test

* Set UT timeout to 40

* Add UT OOMExceptionTest

* Set UT timeout to 50

42910361

C
Fix reduce_sum dtype dispatch bug on gpu (#39349) · 4d7ad277
由 Chen Weihang 提交于 2月 08, 2022
```
* fix pten reduce dispatch bug

* add cast beforce reduce

* fix test failed
```
4d7ad277
L

[bf16] change bf16 print behavior (#39370) · 96964ff8
由 Leo Chen 提交于 2月 08, 2022

96964ff8

07 2月, 2022 1 次提交
- C
  [CustomOp] Support output as input argument of kernel func (#39353) · f1f74e9e
  由 Chen Weihang 提交于 2月 07, 2022
```
* refactor custom op kernel func and utils

* add output sync

* adapte tensor* in utils

* fix windows symbol error
```
  f1f74e9e
06 2月, 2022 1 次提交
- W
  
  [PTEN] Add Gpu context (#39305) · a821c4a9
  由 Wilber 提交于 2月 06, 2022
  
  a821c4a9
04 2月, 2022 2 次提交
- Z
  【Pten】Support data transform in C++ API (#39263) · dcff7fa8
  由 zyfncg 提交于 2月 04, 2022
```
* add data_transform in pten api

* support GetKernelTypeForVar

* fix complie problem of bfloat16

* change error namespace

* add complex type transform unittest

* fix merge conflict
```
  dcff7fa8
- C
  
  remove unchanged infermeta new (#39343) · 0dccdee0
  由 Chen Weihang 提交于 2月 04, 2022
  
  0dccdee0
02 2月, 2022 1 次提交

[PTen] Remove kernel alias name (#39321) · 5dc20c27

由 Chen Weihang 提交于 2月 02, 2022

* remove kernel alias name

* fix depreacted error

* fix deprecated failed

* fix mean error

* resolve conflict

* fix windows failed

5dc20c27

30 1月, 2022 4 次提交

Add a Sparse OP:sparse_csr_to_coo (#39266) · bafea65c

由 zhangkaihuo 提交于 1月 30, 2022

* dense_to_sparse_coo

* optimize unit testing; support rocm

* 1. delete fluid related header file
2. update the copyright

* fix hipMemcpy

* update dense_to_sparsecoo

* add namespace sparse

* sparse_csr_to_dense

* test to_sparse_coo: csr_to_coo

* fix writing error

bafea65c

[PTen] Change all InferMeta functions (#39222) · 7e29cea9

由 Chen Weihang 提交于 1月 30, 2022

* change unary infermeta

* change other infermeta

* change all infermeta format

* resolve conflit

* fix test failed

* resolve reshape conflit

* fix compile failed

* adapt auto api gen

* fix reshape failed

* fix concat failed

* resolve conflict

7e29cea9

Add a Sparse OP : to_sparse_coo (#39264) · 78132fe1

由 zhangkaihuo 提交于 1月 30, 2022

* dense_to_sparse_coo

* optimize unit testing; support rocm

* 1. delete fluid related header file
2. update the copyright

* fix hipMemcpy

* update dense_to_sparsecoo

* add namespace sparse

78132fe1

[pten] fit get all register op kernels (#39288) · eefe5feb

由 Leo Chen 提交于 1月 30, 2022

* upgrade _get_all_register_op_kernels

* add ut

* support xpu/npu

* fix device id

* enhance TransToFluidPlace

* fix compile

eefe5feb

29 1月, 2022 2 次提交

C

rename utils to manual (#39320) · 96bcf2df
由 Chen Weihang 提交于 1月 29, 2022

96bcf2df

[PTen] Tidy pten core headers (#39188) · dd990981

由 Chen Weihang 提交于 1月 29, 2022

* open header for custom kernel

* add core utils

* tidy core code

* tify header

* tidy include

* tidy namespace

* resolve conflit

* fix unittest and coverage

* remove platform using

* resolve conflict

* resolve conflict

* fix digamma namespace error

* fix xpu full kernel error

* fix xpu full kernel error

* polish details

* add place for lib storage

dd990981

28 1月, 2022 5 次提交

C
[PTen] Update all forward argument maping fns (#39252) · 75923a32
由 Chen Weihang 提交于 1月 28, 2022
```
* update forward argument mapping

* fix compile failed

* fix test failed
```
75923a32
Y
[PTen]Refactor scale kernel that has selected_rows input (#39278) · abfc2fe9
由 YuanRisheng 提交于 1月 28, 2022
```
* refactor scale kernel that its input is selected_rows

* complement upload file
```
abfc2fe9

Move digamma to pten (#39240) · 848ae7dc

由 hong 提交于 1月 28, 2022

* move digamma to pten; test=develop

* fix mutable_data bugs; test=develop

* remove useless code; test=develop

* remove kernel compute; test=develop

* fix bug; test=develop

848ae7dc

【Pten】Remove WriteBackOutput in tensor_utils (#39291) · 3ef2922b

由 zyfncg 提交于 1月 28, 2022

* remove remake densetensor

* fix eager test error

* fix bug in eager

* implement AllocateFrom

* remove WriteBackOutput

* fix problem of eager
Co-authored-by: Nzkh2016 <zhangkaihuo@baidu.com>

3ef2922b

Z

Auto-geneate kernel signature in C++ API (#39281) · fc5fa0de
由 zyfncg 提交于 1月 28, 2022

fc5fa0de

27 1月, 2022 11 次提交

Z

implement AllocateFrom (#39280) · d89f246c
由 zhangkaihuo 提交于 1月 27, 2022

d89f246c
C
Add kernelsignature constructor for windows (#39253) · 33e3f5ac
由 Chen Weihang 提交于 1月 27, 2022
```
* add constructor for win

* change impl

* fix bug
```
33e3f5ac
Z
【PTen】Remove ReMakePtenDenseTensor (#39094) · 98c1829b
由 zyfncg 提交于 1月 27, 2022
```
* remove remake densetensor

* fix eager test error

* fix bug in eager
```
98c1829b
Y

refactor elementwise sub grad (#39225) · 7a1e1193
由 YuanRisheng 提交于 1月 27, 2022

7a1e1193

[PTen]Support AllocateFrom in Tensor and Alloc/HostAlloc in Context (#39022) · 5631da9c

由 Aurelius84 提交于 1月 27, 2022

* Support allocate_from in Tensor and allocate_data in Context

* fix #ifdef CUDA

* fix cycle depends

* fix test_xxx_dev_api failed

* fix windows compiling error

* fix unittest

* modify into PImpl

* fix selected rows

* add TODO comment

* refine interface according reviewer

5631da9c

C
[PTen] Add infermeta registry (#39204) · f3f16126
由 Chen Weihang 提交于 1月 27, 2022
```
* add infermeta registry

* add infermeta registry

* add unittest

* polish details
```
f3f16126

[PluggableDevice] Add custom kernel support based on pten kernel management (#38848) · a8879215

由 Aganlengzi 提交于 1月 27, 2022

* [Demo] custom kernel based on pten kernel

* merge and npu custom work well

* del comments

* delete other code

* fix CUDAContext

* fix not found small_vector.h

* support NPU

* fix NPUContext

* fix DeviceContext support

* add UT

* fix call

* add UT

* fix

* fix for comments and ut

* add MACRO control

* fix multi input output

* support env CUSTOM_DEVICE_ROOT

* deal with special cases

* fix for Windows

* try coverage with test_custom_kernel_dot.py

* fix test_custom_kernel_dot

* fix test_custom_kernel_dot

* fix merge

* fix merge

* fix CI

* update

* merge and fix

* remove WITH_CUSTOM_KERNEL

* fix merge

* merge and fix

* fix ut

* fix ut for mac

* add more UT

* add more UT

* fix

a8879215

[pten] add full xpu kernel (#39172) · 93839717

由 chentianyu03 提交于 1月 27, 2022

* add full_kernel xpu

* fix full xpu register device type error

* fix full kernel bug

* add fulllike kernel impl and replace with raw kernel

* fix dev_ctx convert template args error

* modify namespace and header file

* add isinf check

* fix input type args in TensorSetConstantXPU error

93839717

optimize kunlun/xpu softmax_with_cross_entropy add add unitest (#39180) · 2b9bb8bb

由 QingshuChen 提交于 1月 27, 2022

* optimize kunlun/xpu softmax_with_cross_entropy add add unitest
*test=kunlun

* minor
*test=kunlun

* minor
*test=kunlun

* minor
*test=kunlun

* minor
*test=kunlun

2b9bb8bb

Add SparseCooTensor and SparseCsrTensor (#38906) · a7edb3f3

由 zhangkaihuo 提交于 1月 27, 2022

* fix bug:
1. atten: set the default value of attn_dropout_rate to None
2. ffn: add activation parameter

* for pure fp16

* Add a SparseCsrTensor

* remove unused functional

* remove const

* remove SetMemoberTensor

* remove non_zero_nums_, the number of non zero elements of each batch can be obtained from the crows

* SparseCooTensor

* add SetMember

* merge upstream; add SetMember

* merge upstream

* merge upstream; add newline at end of file

* add newline at end of file

* remove newline at end of file

* remove newline at end of file

* stash

* user pten::framework::make_ddim

* user pten::framework::make_ddim

* merge upstream; use the latest mutable_data

* merge upstream; use the latest mutable_data

* return mutable dense tensor

a7edb3f3

F

move math_cuda_utils.h to pten/kernels/funcs (#39246) · 809a10b6
由 Feiyu Chan 提交于 1月 27, 2022

809a10b6

BaiXuePrincess / Paddle 与 Fork 源项目一致

BaiXuePrincess / Paddle
与 Fork 源项目一致