提交 · 57b2033be6d255a250791aff1bc71a7cc429728d · 机器未来 / Paddle

25 1月, 2022 9 次提交

[inference] update trt convert reduce op&ut,test=develop (#39088) · 80753755

由 Zhang Jun 提交于 1月 25, 2022

* [inference] update convert reduce op&ut,test=develop

* update

* update

* update

* add int32 support

* add int32 support

* add comments

* trt < 7.0 do not support int32

* test=develop

* update

* test=develop

80753755

J
[MLU]add mlu kernel for fill_constant op (#39069) · 6e871dbc
由 joeqiao12 提交于 1月 25, 2022
```
* [MLU]add mlu kernel for fill_constant op

* delete device_context DEPS
```
6e871dbc
N
Revert "Replace EigenBroadcast with ElementwiseBroadcast in ReduceGrad (#38959)" (#39205) · 978558be
由 niuliling123 提交于 1月 25, 2022
```
This reverts commit 9059ef69.
```
978558be
J
[MLU]add mlu kernel for split and concat (#39020) · ac3dc0bb
由 joeqiao12 提交于 1月 25, 2022
```
* [MLU]add mlu kernel for concat and split op

* delete device_context DEPS
```
ac3dc0bb
N

Replace EigenBroadcast with ElementwiseBroadcast in ReduceGrad (#38959) · 9059ef69
由 niuliling123 提交于 1月 25, 2022

9059ef69

Optimize nearest_interp forward (#38528) · 232bbce2

由 Lijunhui 提交于 1月 25, 2022

* init commit

* remove comments

* remove nchw branch

* optimize code

* apply fast div mod in 1D kernel, rm 3D kernel

* move init of FastDivMode to CPU

* 3D kernel for nchw, FastDiv for 1D kernel

* debug done. process boundary

* 2^n

* optimize

* optimize

* change code & optimize code

232bbce2

[Move selected_rows PR ] Change the relationship of [include/Cmake]. (#39128) · 2bafd338

由 Weilong Wu 提交于 1月 25, 2022

* Added selected_rows and rw_lock to pten

* Renamed the unit test target to fix CI

* Removed Class SelectedRows in Fluid, changed include/cmake relationship, use pten::SelectedRows in Fluid

* Remove rw_lock.h,rw_lock_test.cc in fluid

* Use pten::RWLock and pten::AutoRDLock, fix CI

* Use pten::SelectedRows

* Use pten::SelectedRows

* Fix to pass NPU CI

* Use pten::SelectedRows, to pass NPU CI

* To fix NPU CI

* To fix NPU CI again

2bafd338

N

[pnorm] fix bug in fp16 & optimize memory (#39011) · 3825b40f
由 Noel 提交于 1月 25, 2022

3825b40f
W

[PTEN] Add xpu context. (#39098) · c1e5a393
由 Wilber 提交于 1月 25, 2022

c1e5a393

24 1月, 2022 7 次提交

[pten] add Scale xpu kernel (#39092) · 7874d0a5

由 chentianyu03 提交于 1月 24, 2022

* add scale xpu kernel

* add scale xpu kernel

* add scale xpu kernel

* replace with pten scale kernel

* change dev_ctx

* modify float16 head path

* remove unused xpu header

7874d0a5

[Pten]Refactor elementwise_add grad / double grad / triple grad Kernel and... · 3bf3a6ee

由 YuanRisheng 提交于 1月 24, 2022

[Pten]Refactor elementwise_add grad / double grad / triple grad Kernel and move them to pten (#39048)

* refactor elementwise add grad

* fix compile bugs

* fix unit test bugs

* fix file conflicts

* fix bugs when buildPtenContext

3bf3a6ee

Remved redundant defintions of likely/unlikely (#38911) · 43919d0a

由 Jacek Czaja 提交于 1月 24, 2022

* - more unlikely

* - compilation fix

* - removed redundant definition

* - fix

* - Fixes

* - compilation fix for windows

43919d0a

[Pten] Migration of eigen numeric extensions and functors in paddle/fluid/operatos/eigen (#39124) · a1e40dc6

由 Feiyu Chan 提交于 1月 24, 2022

* migration of functors in paddle/fluid/operators/eigen and paddle/fluid/platform/eigen_ext.h
* update path of data types like float16.h in includes in extensions.h

a1e40dc6

Z

unify compare functor (#39024) · def81b4f
由 Zhang Ting 提交于 1月 24, 2022

def81b4f

[PTEN] Move dynload from fluid to pten. (#39120) · 3c1dc6f6

由 Wilber 提交于 1月 24, 2022

* move dynload from fluid to pten.

* fix ci compile

* fix windows ci compile.

* update

* update

* fix compile error

3c1dc6f6

support sparse of adam, *test=kunlun (#38483) · e106901e

由 z8hanghuan 提交于 1月 24, 2022

* support sparse of adam, *test=kunlun

* add pre-commit-config.yaml

* support sparse of adam in KL2,*test=kunlun

* support sparse of adam in KL2, *test=kunlun

* modify xpu.cmake, *test=kunlun

* support sparse of adam, rm some wait, *test=kunlun

* support sparse of adam, rm some wait, *test=kunlun

* support sparse of adam, *test=kunlun

* support sparse of adam, *test=kunlun

* support sparse of adam, *test=kunlun

* support sparse of adam, *test=kunlun

* support sparse of adam, *test=kunlun

e106901e

21 1月, 2022 12 次提交
- C
  [pten] fix test concat dev api build failed (#39117) · a14dc688
  由 chentianyu03 提交于 1月 21, 2022
```
* fix test concat dev api build failed

* fix conflict

* fix conflict
```
  a14dc688
- Y
  [PTen]Separate origin Kernel and add Kernel for C++ API (#39002) · a0f586bc
  由 YuanRisheng 提交于 1月 21, 2022
```
* add kernel for c++ api

* fix compile bugs

* fix kunlun compile bugs

* perfect cmake

* fix compile bugs when run ci-inference

* fix compile bugs

* add non-raw kernel for fluid op

* fix compile bugs

* fix compile bugs

* fix unit test bug
```
  a0f586bc
- C
  
  [pten] add concat pten kernel (#38955) · 06803c29
  由 chentianyu03 提交于 1月 21, 2022
  
  06803c29
- W
  
  Renamed selected_rows.* -> selected_rows_utils.* (#39037) · 814e5ab4
  由 Weilong Wu 提交于 1月 21, 2022
  
  814e5ab4
- Z
  
  modify DivideFunctor to match ElementwiseSameDims template (#39041) · df515255
  由 Zhang Ting 提交于 1月 21, 2022
  
  df515255
- T
  Keep strided_slice op behavior consistent with slice op when starts input is... · b47fb764
  由 TeslaZhao 提交于 1月 21, 2022
```
Keep strided_slice op behavior consistent with slice op when starts input is less than -rank (#39066)
```
  b47fb764
- F
  [MLU]add mlu ci dockerfile (#39021) · fdab43b5
  由 fwenguang 提交于 1月 21, 2022
```
* [MLU]add mlu ci dockerfile

* fix comment

* add cncl
```
  fdab43b5
- A
  [PTen]Migrate Dim and DDim from paddle::framework into pten namespace (#39053) · 4e23ba32
  由 Aurelius84 提交于 1月 21, 2022
```
* Migrate Dim and DDim from paddle::framework into pten namespace

* fix paddle::framework::Array

* fix framework::Array
```
  4e23ba32
- R
  
  fix npu c_allgather int64 (#39099) · 89f903da
  由 ronnywang 提交于 1月 21, 2022
  
  89f903da
- F
  add block and grid loop for index_sample kernel to deal with a large-shape tensor (#37816) · 4adeff06
  由 FlyingQianMM 提交于 1月 21, 2022
```
* add block and grid loop for index_sample kernel to deal with a large-shape tensor

* fix code format

* limit grid dim
```
  4adeff06
- F
  
  [MLU]add batch_norm mlu kernel (#39070) · 29796efe
  由 fwenguang 提交于 1月 21, 2022
  
  29796efe
- W
  [PTEN] Add cpu context (#38979) · 064bc4b8
  由 Wilber 提交于 1月 21, 2022
```
* add cpu_context.

* update

* update

* update

* update

* update

* fix ci problem

* fix npu ci problem

* update

* fix ci compile
```
  064bc4b8
20 1月, 2022 7 次提交
- F
  
  [MLU]add mlu kernel for top_k and top_k_v2 (#39065) · e02dec01
  由 fwenguang 提交于 1月 20, 2022
  
  e02dec01
- F
  
  [MLU]add mlu kernel for cast and scale op (#38961) · e3e50ea8
  由 fwenguang 提交于 1月 20, 2022
  
  e3e50ea8
- A
  [Pten] Migrate bfloat16/float16/complex from paddle::platform into pten::common (#39044) · f1143f0c
  由 Aurelius84 提交于 1月 20, 2022
```
* Migrate bfloat16/float16/complex from platform into pten::common

* fix typo

* fix code style
```
  f1143f0c
- Y
  
  mod communicator (#39064) · 2a9c993e
  由 yaoxuefeng 提交于 1月 20, 2022
  
  2a9c993e
- Z
  Fix master weight bug for multi_tensor optimizer(momentum, adam) (#38991) · 6b0c57cf
  由 zhangbo9674 提交于 1月 20, 2022
```
* fix mp

* support merged_momentum for mp
```
  6b0c57cf
- S
  
  remove if !defined(WIN32) (#39058) · 90e9233a
  由 sneaxiy 提交于 1月 20, 2022
  
  90e9233a
- S
  
  fix gelu compile on CUDA 10 (#39045) · 0617a3ed
  由 sneaxiy 提交于 1月 20, 2022
  
  0617a3ed
19 1月, 2022 1 次提交
- Z
  
  Add conv2d_transpose and conv2d_transpose_grad for XPU,test=kunlun (#38956) · c7de7440
  由 zhangyikun02 提交于 1月 19, 2022
  
  c7de7440
18 1月, 2022 4 次提交

Mish FP32/BF16 kernel, conv and fc fuse passes (#38623) · 1d18bc2c

由 Sławomir Siwek 提交于 1月 18, 2022

* Mish

* Change exp() library

* mish fuse pass

* mish attrs

* fixes

* mishop maker

* remove attrs

* mish kernal for bf16

* fc+mish fuse

* fix code format error

* Resolve merge conflicts

* Update mish operator version

* update mish variable to new naming convention

1d18bc2c

change CUDA implementaion of uniform/gaussian OP (#38611) · bbbd75e4
由 zhouweiwei2014 提交于 1月 18, 2022
```
* change CUDA implementaion of uniform/gaussian OP

* fix unittest
```
bbbd75e4

[Unify Tensors PR #8] Merged Tensor into DenseTensor, test=allcases (#38914) · 2052f1e3

由 Zhanlue Yang 提交于 1月 18, 2022

* Merged LoDTensor with Tensor,test=allcases

* Patched python level LoDTensor

* Patched python level LoDTensor

* Merge Tensor into DenseTensor

* Fixed namespace issues,test=allcases

* Fixed merge issues

* Fixed inference issues

* Fixed NPU test issues

* Fixed merge issues

2052f1e3

T

fix lookup_table_v2 error in kunlun2 (#38855) · df898f8b
由 taixiurong 提交于 1月 18, 2022

df898f8b

机器未来 / Paddle 与 Fork 源项目一致

机器未来 / Paddle
与 Fork 源项目一致