提交 · a863b32ea8441b2448e487b417e7ce596f530a44 · BaiXuePrincess / Paddle

21 2月, 2022 7 次提交
- 0
  [Dy2St]Fix cond grad error when handle tensor array (#39689) · a863b32e
  由 0x45f 提交于 2月 21, 2022
```
* fix cond grad error when handle tensor array

* add UT
```
  a863b32e
- C
  [pten]rm reduce_sum and reduce_mean raw kernel (#39484) · 2bb5aae8
  由 chentianyu03 提交于 2月 21, 2022
```
* rm reduce_sum raw kernel

* remove reduce_mean kernel

* remove reduce_mean kernel

* reduce support int and int64_t

* mean support int and int64_t type
```
  2bb5aae8
- Z
  [bf16] add bf16 kernel: elementwise_max (#39461) · 93016331
  由 zhangbo9674 提交于 2月 21, 2022
```
* add elementwise_max & unittest

* refine cuda register and unittest

* refine unittest

* refine uinttest for bf16

* refine optest

* refine code

* refine unittest

* refine unittest
```
  93016331
- Y
  [PTen]Remove infershape of Reshape OP (#39631) · 45dd4a5f
  由 YuanRisheng 提交于 2月 21, 2022
```
* remove infershape and Xshape

* add xshape

* fix bugs when run ci

* fix bugs when run ci

* fix bugs when run infrt test

* pass converage
```
  45dd4a5f
- Z
  [HeterPS]fix ut for heteps comm op (#39684) · d41836ef
  由 zmxdream 提交于 2月 21, 2022
```
* fix. test=develop

* fix. test=develop

* fix code style. test=develop

* fix. test=develop

* fix. test=develop
```
  d41836ef
- S
  
  fix alignment bug (#39747) · 65ced1fa
  由 sneaxiy 提交于 2月 21, 2022
  
  65ced1fa
- S
  
  fix bug: core when missing range XPU kernel in kunlun2 (#39673) · 496aadfb
  由 ShiningZhang 提交于 2月 21, 2022
  
  496aadfb
20 2月, 2022 4 次提交

[PTen->Phi PR1] Change pten dirname and namespace to phi (#39748) · dcfe1986

由 Chen Weihang 提交于 2月 20, 2022

* rename pten dir to phi

* rename namespace to phi

* rename infrt pten dir to phi

* resolve conflict

* rename pten to phi in cmake

* revert all infrt change

* change needed files

* fix infrt failed

* fix inference failed

dcfe1986

add index initialization in the block loop for index_sample kernel when... · c6950ab2

由 FlyingQianMM 提交于 2月 20, 2022

add index initialization in the block loop for index_sample kernel when dealing with a input tensor whose shape is larger than block_dim * grid_dim (#39736)

* add block and grid loop for index_sample kernel to deal with a large-shape tensor

* fix code format

* limit grid dim

* fix the omissive initialization of index_i in the second cycle for index_sample kernel

* fix conflicts

c6950ab2

Y

Rename the general elementwise and broadcast functions. (#39623) · 553afc07
由 Yiqun Liu 提交于 2月 20, 2022

553afc07
S
Add int16 support for several ops (#39636) · 267275d9
由 sneaxiy 提交于 2月 20, 2022
```
* add more op int16 support

* fix xpu ci
```
267275d9

19 2月, 2022 5 次提交

[Pten]Unify paddle/pten::framework::ddim into pten::ddim (#39614) · 2fe04264

由 Aurelius84 提交于 2月 19, 2022

* Unify paddle/pten::framework::ddim into pten::ddim

* fix paddle namespace

* compile sucessfully

* fix npu src file

* fix conflict

* fix conflict

* fix tensorrt compiler error

* fix conflict

* fix conflict

* fix tesst file conflict

* fix conflict

* fix mlu file conflict

* fix mlu file conflict

* fix cinn header file conflict

* fix conflict

* fix conflict

* fix conflict

* fix conflict

2fe04264

[Pten] Add selected_rows kernel for Full (#39465) · 79f8eeca

由 zyfncg 提交于 2月 19, 2022

* Add selected_rows kernel for full

* remove fill_constant register in fluid

* fix bug without GPU

* add jit_kernel_helper dependency for fc

* do some refactor

* add unittest for ops signatures

* add coverage unittest

* fix merge conflict

* fix full selectew_rows bug

79f8eeca

fix RecordEvent interface (#39675) · 019a552b

由 chenjian 提交于 2月 19, 2022

* fix RecordEvent interface

* modify default level to 4

* update interface use

* add const default trace level

* update operator.cc

019a552b

[Pten] Adjust the params of creation kernel for inference (#39573) · 4e5d6743

由 zyfncg 提交于 2月 19, 2022

* remove manual_api

* change sig map of full and empty

* fix fill_any_like_xpu_op

* fix fill_any_like_xpu_op

* fix problem of fill_any_like_xpu_op

* fix conflict

* polish code

4e5d6743

Add the DistributedFusedLamb optimizer (#39148) · 5df3cd61

由 sneaxiy 提交于 2月 19, 2022

* add DistributedFusedLamb op

* polish code

* fix compile error

* compatible with pten changement

* fix rocm compile error

* improve converage

* update upstream/develop

* fix cast_with_ptr.h

* add FLAGS_distributed_lamb_divide_nranks_when_allreduce=1

* fix clip before allreduce

* add use_master_param_norm

* code polish

* fix bug

* fix ROCM ci

5df3cd61

18 2月, 2022 9 次提交
- F
  [Pten] blas and lapck migration (#39587) · 8c7ee8c2
  由 Feiyu Chan 提交于 2月 18, 2022
```
* move blas related files
* move lapack related files
```
  8c7ee8c2
- T
  cinn_instruction_run_op test (#39576) · fdc4fe3b
  由 TeFeng Chen 提交于 2月 18, 2022
```
* add cinn_instruction_run_op test code

* update several interfaces of CinnLaunchContext

* update several interfaces and add detail comments in CinnLaunchContext class

* to skip the bug of error message check

* fix ut test failed due to reliant interface updated
```
  fdc4fe3b
- X
  [pten] trans diagonal kernel into pten (#39575) · 5c66338f
  由 xiongkun 提交于 2月 18, 2022
```
* trans diagonal kernel into pten

* fix by code review
```
  5c66338f
- A
  [IPU] Update IpuStrategy (#39644) · 46161679
  由 Allen Guo 提交于 2月 18, 2022
```
* Update IpuStrategy

* fix ci

* rerun ci
```
  46161679
- B
  
  refactor the forward implementation of shape npu op (#39613) · e674af23
  由 baoachun 提交于 2月 18, 2022
  
  e674af23
- T
  
  dropout support Seed, fix elementwise_add_grad bug, test=kunlun (#39656) · 70b9f2ac
  由 taixiurong 提交于 2月 18, 2022
  
  70b9f2ac
- Q
  [MLU]add matmul and matmul_v2 op (#39539) · 229ec32a
  由 qipengh 提交于 2月 18, 2022
```
* [MLU]add matmul and matmul_v2 op

* [MLU] fix data_type and del matmul

* [MLU] fix compile error

* [MLU] fix ci_check error
```
  229ec32a
- J
  
  add flatten op for mlu (#39530) · 4c5cec5c
  由 joeqiao12 提交于 2月 18, 2022
  
  4c5cec5c
- Z
  [MLU]add sync stream ops and broadcast pytest (#39518) · d2bd05b9
  由 zn 提交于 2月 18, 2022
```
* [MLU]add sync stream ops and broadcast pytest

* [MLU]fix broadcast pytest to add data type
```
  d2bd05b9
17 2月, 2022 7 次提交
- L
  [pten] move bernoulli kernel to pten (#39590) · f86073c4
  由 Leo Chen 提交于 2月 17, 2022
```
* move bernoulli kernel to pten

* follow comments
```
  f86073c4
- J
  
  add reshape2 op for mlu (#39562) · 2d2f11d1
  由 joeqiao12 提交于 2月 17, 2022
  
  2d2f11d1
- S
  move trunc to pten (#39543) · 4501abd6
  由 Sing_chan 提交于 2月 17, 2022
```
* move trunc to pten

* modify according to YuanRisheng's comment
```
  4501abd6
- H
  add softplus op for kunlun2. test=kunlun (#39555) · 9f99b591
  由 houj04 提交于 2月 17, 2022
```
* add softplus op for kunlun2. test=kunlun

* add softplus op for kunlun2. test=kunlun

* fix code style. test=kunlun

* fix code style. test=kunlun

* add more test cases. test=kunlun
```
  9f99b591
- Z
  [Pten] Remove register of matmul_v2 kernel (#39542) · db43b541
  由 zyfncg 提交于 2月 17, 2022
```
* remove register of matmul_v2 kernel

* delete matmul_v2 grad register in fluid
```
  db43b541
- C
  
  move trace infer shape (#39517) · 1c9b2483
  由 Chen Weihang 提交于 2月 17, 2022
  
  1c9b2483
- N
  
  Modified distribution kernel with Kernel Primitive API (#39563) · 1354652b
  由 niuliling123 提交于 2月 17, 2022
  
  1354652b
16 2月, 2022 8 次提交
- T
  
  optimize prior_box for kunlun, *test=kunlun (#39477) · e254e7c6
  由 TTerror 提交于 2月 16, 2022
  
  e254e7c6
- F
  
  [MLU] support adative pooling (#39500) · f138371c
  由 fwenguang 提交于 2月 16, 2022
  
  f138371c
- 0
  Move lerp OP to pten (#39524) · d480d7b1
  由 0x45f 提交于 2月 16, 2022
```
* move lerp to pten

* refine include

* move files

* refine code
```
  d480d7b1
- A
  
  Add ConditionalBlockGradInferVarType (#39585) · ff7e3590
  由 Aurelius84 提交于 2月 16, 2022
  
  ff7e3590
- L
  [bf16] pten matmul cuda kernel support bf16 (#39485) · d5a0d31a
  由 Leo Chen 提交于 2月 16, 2022
```
* pten matmul cuda kernel support bf16

* fix pten kernel name

* add matmul_grad bf16 kernel

* add emptylike bf16 kernel

* fix compile

* suppport rocm

* fix error

* fix rocm

* add bf16 header file

* fix compile
```
  d5a0d31a
- F
  [Pten] move complex_functors.h (#39558) · 5b5656d0
  由 Feiyu Chan 提交于 2月 16, 2022
```
* move complex_functors.h and update all references to symbols within it
```
  5b5656d0
- C
  [PTen] Rename general grad infermeta func (#39578) · 12ca438e
  由 Chen Weihang 提交于 2月 16, 2022
```
* rename general grad infermeta func

* remove useless code
```
  12ca438e
- A
  [Pten]Modify framework::VisitDataType into Pten::VisitDataType (#39550) · 6b756fb7
  由 Aurelius84 提交于 2月 16, 2022
```
* Modify framework::VisitDataType into Pten::VisitDataType

* migrate unittest
```
  6b756fb7

BaiXuePrincess / Paddle 与 Fork 源项目一致

BaiXuePrincess / Paddle
与 Fork 源项目一致