提交 · 8b1048b460e6f7093746e648db39604200c7c5cd · PaddlePaddle / Paddle

16 2月, 2022 10 次提交
- L
  Revert "[pten] remove concat fluid kernel (#39268)" · 8b1048b4
  由 Leo Chen 提交于 2月 16, 2022
```
This reverts commit 552db8dc.
```
  8b1048b4
- T
  
  optimize prior_box for kunlun, *test=kunlun (#39477) · e254e7c6
  由 TTerror 提交于 2月 16, 2022
  
  e254e7c6
- F
  
  [MLU] support adative pooling (#39500) · f138371c
  由 fwenguang 提交于 2月 16, 2022
  
  f138371c
- 0
  Move lerp OP to pten (#39524) · d480d7b1
  由 0x45f 提交于 2月 16, 2022
```
* move lerp to pten

* refine include

* move files

* refine code
```
  d480d7b1
- A
  
  Add ConditionalBlockGradInferVarType (#39585) · ff7e3590
  由 Aurelius84 提交于 2月 16, 2022
  
  ff7e3590
- L
  [bf16] pten matmul cuda kernel support bf16 (#39485) · d5a0d31a
  由 Leo Chen 提交于 2月 16, 2022
```
* pten matmul cuda kernel support bf16

* fix pten kernel name

* add matmul_grad bf16 kernel

* add emptylike bf16 kernel

* fix compile

* suppport rocm

* fix error

* fix rocm

* add bf16 header file

* fix compile
```
  d5a0d31a
- F
  [Pten] move complex_functors.h (#39558) · 5b5656d0
  由 Feiyu Chan 提交于 2月 16, 2022
```
* move complex_functors.h and update all references to symbols within it
```
  5b5656d0
- C
  [PTen] Rename general grad infermeta func (#39578) · 12ca438e
  由 Chen Weihang 提交于 2月 16, 2022
```
* rename general grad infermeta func

* remove useless code
```
  12ca438e
- A
  [Pten]Modify framework::VisitDataType into Pten::VisitDataType (#39550) · 6b756fb7
  由 Aurelius84 提交于 2月 16, 2022
```
* Modify framework::VisitDataType into Pten::VisitDataType

* migrate unittest
```
  6b756fb7
- Y
  [Pten]Remove reshape and elementwise_add's registry code in Fluid (#39317) · c6478270
  由 YuanRisheng 提交于 2月 16, 2022
```
* remove reshape and elementwise_add registry

* delete code

* fix bugs when run ci ut

* remove log

* fix bugs when run unit test

* fix bugs when run unit test

* fix bugs when run cinn

* fix bugs when run ci-mac-python3

* fix compile bugs

* fix compile bugs

* fix compile bugs

* fix bugs when run kunlun

* fix bugs when compile

* update code according comment
```
  c6478270
15 2月, 2022 10 次提交

J

disabled unnecessary int reorders profiling (#39498) · 3581c075
由 jakpiase 提交于 2月 15, 2022

3581c075

[PluggableDevice] Add custom runtime support (#38740) · 3e7825f3

由 ronnywang 提交于 2月 15, 2022

* [CustomRuntime] Add DeviceManager

* [CustomRuntime] Add DeviceInterface

* [CustomRuntime] Add Stream, Event, DeviceGuard, CallbackManager

* [CustomRuntime] Add plug-in device

* [CustomRuntime] Memory module support PluggableDevice

* [CustomRuntime] Add WITH_PLUGGABLE_DEVICE cmake option

* update

* [API] update API doc based on comments, test=develop
Co-authored-by: Nqili93 <qili93@qq.com>

3e7825f3

F
[Pten] move paddle/operators/math/functors.h and compound_functors.h (#39514) · 0d46a108
由 Feiyu Chan 提交于 2月 15, 2022
```
* move paddle/operators/math/functors.h
* move paddle/operators/math/compound_functors.h
```
0d46a108

Add cinn_instruction_run_op for launching execution of a cinn instruction (#39435) · 9d0baeab

由 TeFeng Chen 提交于 2月 15, 2022

* add cinn_instruction_run_op for launching execution of a cinn instruction

* fix multi definition compilation error

* update cmake

* fix bug at infershape

* fix compile error due to lacking header file

9d0baeab

move histogram to pten (#39496) · 556f6eb0

由 hong 提交于 2月 15, 2022

* move histogram to pten; test=develop

* fix format error; test=develop

* fix histogram kernel format; test=develop

556f6eb0

Move Abs OP to pten (#39492) · fb473067

由 From00 提交于 2月 15, 2022

* Move Abs op to pten

* Fix NPU compilation error

* Fix CI error

* Use LaunchSameDimsElementwiseCudaKernel in pten

fb473067

S

add dropout fp32 (#39501) · b81358d1
由 sneaxiy 提交于 2月 15, 2022

b81358d1

move algorithm.h (#39502) · 7eb9593e

由 Feiyu Chan 提交于 2月 15, 2022

Move paddle/fluid/operators/math/algorithm.h to paddle/pten/kernels/funcs and rename all references to symbols in it.

7eb9593e

[Pten]Move expand_v2 to pten (#39471) · 2d16d69b

由 Linjie Chen 提交于 2月 15, 2022

* move expand to pten

* move expand_v2 to pten

* move expand_v2 to pten

* fix grad register

* fix grad register

* fix tensorcpry

* fix tensorcopy

* fix tensorcopy

* fix tensorcopy

* fix tensorcopy

* fix ci

* fix tensorcopy

2d16d69b

[PTen]Migrate proto::VarType outside of Pten (#39411) · 7e7e9404

由 Aurelius84 提交于 2月 15, 2022

* #1 migrate dist-related type()-> dtype()

* move datatype function from pten -> fluid/framework

* change type() in imperative into convert(dtype())

* modify xx_tensor->type into xx_tensor->dtype

* change the set_type interface and the caller

* modify xx_tensor.type into xx_tensor.dtype

* fix mutable_data(place, dtype())

* change caller of mutable_data in pten and distributed

* change the caller of mutable_data in fluid/framework

* change the caller of mutable_data in imperative directory

* mutable_data: inference

* update the call of mutable_data

* transfer MakePenScalarArray MakePtenScalar ResetHolderWithType

* pass the compile. the next step is remove VarType in Pten

* fix all and remove VarType from pten. success in linux. Next task is other platform

* fix conflict with develop

* fix compiled error

* Fix reset conversion

* fix conflict

* fix compiled problem

* fix typo

* Fix << in tensor_utils.cc

* fix type->dtype

* fix unittest

* fix tensor init constructor

* fix DataTypeSize for BFloat16

* fix code style

* fix npu compiled error

* fix npu

* compile npu sucessfully

* fix conflict

* fix conflict
Co-authored-by: Nxiongkun <xiongkun03@baidu.com>

7e7e9404

14 2月, 2022 4 次提交

C
[PTen] Add HasAttr for ArgumentMappingContext (#39464) · ddb1e23f
由 Chen Weihang 提交于 2月 14, 2022
```
* add has_attr for arg map context

* skip useless attr now

* skip attr if not exists

* fix typo
```
ddb1e23f

[pten] add split kernel (#39060) · d0df5632

由 chentianyu03 提交于 2月 14, 2022

* add split kernel

* add split kernel signature

* fix split bug

* modify MakePtenScalarArrayFromVarList

* modify MakePtenScalarArrayFromVarList

* fix split windows register error

* add test case for split kernel

* replace raw split kernel with pten kernel

* fix makeScalar/ScalarArray bug

* remove debug log

* remove int64_t type in buildPtcontext

* update by code review

* fix split dev test failed

* change DenseTensorMeta to MetaTensor

* change split api code from auto gen to manual

* split cuda kernel support bfloat16 type

* fix conflict

* rm raw split kernel

* merge develop branch

* change to pten::errors

d0df5632

T

fix gather_nd, *test=kunlun (#39283) · d12c3636
由 TTerror 提交于 2月 14, 2022

d12c3636
[MLU] add mlu kernel for c_broadcast op (#39470) · 1b9e6790
由 mhhhh1 提交于 2月 14, 2022

1b9e6790

11 2月, 2022 11 次提交
- L
  
  Add TensorRT inspector into Paddle-TRT (#38362) · 69793a27
  由 Leo Chen 提交于 2月 11, 2022
  
  69793a27
- J
  Added shape (U)INT8/BF16/FP32 oneDNN kernel (#36033) · 52bbaae9
  由 jakpiase 提交于 2月 11, 2022
```
* added shape oneDNN kernel

* removed unnecessary import from test

* added skipping tests for GPU

* refactoring

* refactored shape kernel

* added tests in new framework

* removed one line

* minor change

* added newline at EOF

* added formatting

* added attributes as extra
```
  52bbaae9
- J
  
  uniform_random op for mlu (#39450) · 02f06708
  由 joeqiao12 提交于 2月 11, 2022
  
  02f06708
- Z
  [bf16] add bf16 kernel: transpose & unbind (#39457) · 1e6047f1
  由 zhangbo9674 提交于 2月 11, 2022
```
* add transpose unbind

* add unittest

* refine transpose unittest
```
  1e6047f1
- Z
  [MLU]support c_gen_cncl_id_op run on MLU device (#39336) · 89aa8b1a
  由 zn 提交于 2月 11, 2022
```
Co-authored-by: Nzhangna <zhangna@cambricon.com>
```
  89aa8b1a
- F
  
  [MLU] add pool2d and pool2d_grad mlu kernel (#39453) · 702bce57
  由 fwenguang 提交于 2月 11, 2022
  
  702bce57
- F
  [Pten] move operators/math/math_function_* to pten/kernels/func (#39300) · d25a7f9e
  由 Feiyu Chan 提交于 2月 11, 2022
```
* move operators/math/math_function_* to pten/kernels/func
* namespace from `paddle::operators::math` to `pten::funcs`
```
  d25a7f9e
- Z
  Optimize performance of softmax_bwd when axis!=-1 (#38609) · 2ea15fc9
  由 Zhang Zheng 提交于 2月 11, 2022
```
* Optimize performance of softmax_bwd when axis!=-1

* fix

* fix

* fix

* fix
```
  2ea15fc9
- L
  Optimize bilinear interpolation foward (#39243) · a1174973
  由 Lijunhui 提交于 2月 11, 2022
```
* bilinear_fw init

* optimize code

* pre-compute linear_interp input index
```
  a1174973
- C
  [PTen] Move grad GetExpectedPtenKernelArgs into pten (#39418) · 667bd962
  由 Chen Weihang 提交于 2月 11, 2022
```
* move grad get expected pten kernel args

* fix reduce sum error

* fix element_sub_grad failed

* revert kernel judge change
```
  667bd962
- Z
  Support different dtypes of inputs for elementwise ops (#38859) · bf305033
  由 Zhang Ting 提交于 2月 11, 2022
```
* improve backward performance

* support different dtypes for elementwise ops
```
  bf305033
10 2月, 2022 5 次提交

F
[MLU] add mlu kernel for accuracy op (#39337) · 383de295
由 fwenguang 提交于 2月 10, 2022
```
* [MLU] add mlu kernel for accuracy op

* fix license format

* fix error message
```
383de295
F
[NPU] add reduce_min (#39019) · 2b8b16d7
由 furnace 提交于 2月 10, 2022
```
[NPU] add reduce_min
```
2b8b16d7

move Masked select to pten (#39193) · e2ad433b

由 hong 提交于 2月 10, 2022

* move masked select cpu kernel

* add masked selected gpu kernel; test=develop

* fix bugs; test=develop

* bug fix; test=develop

* bug fix; test=develop

* add namespace to set mask array; test=develop

* fix bug; test=develop

* fix bugs; test=develop

* fix ddim bug; test=develop

* fix npu op bug; test=develop

* fix xpu dependecy bug; test=develop

* move kernel args to sig.cc; test=develop

e2ad433b

Modify the unsqueeze dimension of input data in conv1d NCL And NLC format (#38425) · 224bc511

由 crystal 提交于 2月 10, 2022

* optimize conv1d forward

* add conv opt

* Optimize memory copy

* delete share data with

* set num_filters=512

* add nlc optimize

* Optimize num_filter=512 data on A100 and V100

* Fix the workspace_size size setting of filter

224bc511

Z
[bf16] add bf16 kernel: squeeze & unsqueeze & stack (#39402) · 59c7aea5
由 zhangbo9674 提交于 2月 10, 2022
```
* add squeeze unsqueeze stack

* add unittest

* add cpu kernel
```
59c7aea5

PaddlePaddle / Paddle 1 年多 前同步成功

PaddlePaddle / Paddle
1 年多前同步成功