提交 · 23c05f2f0c8991950c3f011a5ad626f05ec486e3 · BaiXuePrincess / Paddle

20 10月, 2022 3 次提交

K
[cherry pick] Add FusedMultiTransformer fuse pass for GPT3 (#47150) · 396427a7
由 Kaipeng Deng 提交于 10月 20, 2022
```
* add fused_attention_pass. test=develop

* support fp16. test=develop

* fix format. test=develop
```
396427a7

[cherry-pick] Fix quantize model deploy bug in MKLDNN (#47119) · c2d344dd

由 yeliang2258 提交于 10月 20, 2022

* Fix quantize model deploy bugs when using MKLDNN (#45920)

* fix immutable op quantize bugs

* fix

* fix build bug

* fix test

* notest,test=inference

* fix ppyoloe acc drop bugs

* fix test

* fix test

* add test

* fix

* fix

* fix test

* fix refined name bug

* fix test

* bias fix

* fix matmul weight dequant bug

* re-ci

* fix tester

* fix test

* fix tester

* update weight dequantize func

* update code

* update test for converage

* update test

* update cmake

* update cmakelist

* update code

* rerun ci

* remove useless code

* re-ci

* update code

* update code

* fix header

* update code for log

c2d344dd

W
[Cherry-pick] layernorm shift partation enhance (#47086) · 9ed1454a
由 Wang Bojun 提交于 10月 20, 2022
```
* Enhance the layernorm shift partation fuse op when shift size > 0 (roll shifting)
* fix cherry-pick test
```
9ed1454a

19 10月, 2022 3 次提交

Add unsigned int8 scale propagation (#46378) (#47156) · 66dccd7d

由 yeliang2258 提交于 10月 19, 2022

* Add unsigned int8 propagation

* Add or modify unit tests

* Correct concat scale checking

* Apply review suggestions

* Corrections
Co-authored-by: Njoanna.wozna.intel <joanna.wozna@intel.com>

66dccd7d

Add enable_partial_send_recv switch in pipeline_configs (#46992) (#47083) · 1d015f12

由 Ghost Screaming 提交于 10月 19, 2022

* Fix bug of reduce_sum op. When input.numel() > INT32_MAX, its result
is wrong.

* Support allow_partial switch, which can be configure in
pipeline_configs. If sent tensor are not the same from
different hosts, they shouldn't been sent partially and
then concated as a whole tensor.

* Change name allow_partial to enable_partial_send_recv.

* Add global variable _enable_partial_send_recv

1d015f12

W
[Dy2St]Fix recurrent op eager deletion pass error in dy2st (#47105) (#47134) · 69515e90
由 WangZhen 提交于 10月 19, 2022
```
[CherryPick][Dy2St]Fix recurrent op eager deletion pass error in dy2st
```
69515e90

17 10月, 2022 2 次提交

Z
[cherry-pick]Sparse static graph (#46838) · 10225d22
由 zhangkaihuo 提交于 10月 17, 2022
```
cherry-pick : #46322, #46245
Sparse API 支持静态图
```
10225d22

[IPU] paddle-inference support custom-ops (#45235) (#46868) · bd89be12

由 Allen Guo 提交于 10月 17, 2022

* paddle-inference support custom-ops
Co-authored-by: NZhixin Yao <zhixiny@graphcore.ai>

* fix tolower
Co-authored-by: NZhixin Yao <zhixiny@graphcore.ai>
Co-authored-by: NZhixin Yao <zhixiny@graphcore.ai>

bd89be12

14 10月, 2022 1 次提交
- Z
  
  [Paddle-TRT] support new quant format from slim (#46022) (#46979) · b8677c0d
  由 zhoutianzi666 提交于 10月 14, 2022
  
  b8677c0d
13 10月, 2022 1 次提交
- Z
  
  interpretercore thread not always spin (#46687) (#46952) · d90aaa6e
  由 zhangbo9674 提交于 10月 13, 2022
  
  d90aaa6e
11 10月, 2022 1 次提交

[cherry-pick] [PHI] relu6_grad kernel (#46501) (#46862) · 2bcbf8b0

由 Sławomir Siwek 提交于 10月 11, 2022

* [PHI] Migrate gelu kernels (#45596)

* gaussian random

* mkldnn to onednn renaming

* fix merge conflicts

* remove fluid code

* onednn renaming

* gelu fwd

* sort activations

* gelu gradient

* remove unused macros

* merge conflicts

* fix merge conflicts

* remove extra contraint from gelu op

* [PHI] relu6_grad kernel (#46501)

* Relu6

* remove fluid handler

* add individual kernel signature

* coding style

* replace bounded_relu with clip

* whitespace

* code style

2bcbf8b0

27 9月, 2022 1 次提交

[cherry-pick] clear extra attrs of some ops in OpMaker (#45845, #45984, 46060) (#46218) · 0cc2251f

由 zyfncg 提交于 9月 27, 2022

* Clear extra attrs of elementwise op in OpMaker (#45845)

* clear extra attrs of elementwise op in opmaker

* fix op_debug_string_test

* fix bug of grad_add

* fix sort of runtime attrs

* Clear extra attrs of scale in OpMaker (#45984)

* clear extra attr of scale in opmaker

* fix sum bug

* fix merge conflict

* fix minus

* Clear extra attributes of some Op in OpMaker (Part4) (#46060)

* clear extra attr of some ops in opmaker

* revert clear use_cudnn for pool

* fix test_operator_desc

* fix Attr interface of OperatorBase

* fix code stype

0cc2251f

23 9月, 2022 1 次提交
- Z
  
  fix compile problem (#46354), test=kunlun (#46383) · 6a508334
  由 zyfncg 提交于 9月 23, 2022
  
  6a508334
20 9月, 2022 5 次提交
- H
  [PolishComments] Polish some code comments (#46032) (#46261) · 42e56f65
  由 HongyuJia 提交于 9月 20, 2022
```
* polish code comments

* polish data_device_transform.cc
```
  42e56f65
- L
  [cherry-pick] Refine thread pool config of interpretercore (#46219) · 1418a719
  由 Leo Chen 提交于 9月 20, 2022
```
* add config

* add config

* follow comments

* fix serial run
```
  1418a719
- Z
  [Inference] fix preln_residual_bias_fuse_pass bug in TNT_small model (#46178) (#46260) · c384b00d
  由 zhoutianzi666 提交于 9月 20, 2022
```
* fix preln_residual_bias_fuse_pass bug in TNT_small model
```
  c384b00d
- Z
  Run_program_op add scope cache & reuse (#45813) (#46223) · 4f28a4c2
  由 zhangbo9674 提交于 9月 20, 2022
```
* add scope cache & reuse

* add gc scope for end of each train step

* del scope reuse for jit

* refine code

* test
```
  4f28a4c2
- Z
  Fix wrong eigen header include (#46082) (#46202) · ac8cce20
  由 zyfncg 提交于 9月 20, 2022
```
* fix wrong eigen header include

* fix complie bug

* fix nan_inf_utils_detail

* fix resource_manager

* fix conv_miopen_helper
```
  ac8cce20
19 9月, 2022 1 次提交
- X
  
  convfusion_cache (#46054) · f4ec1563
  由 xiaoxiaohehe001 提交于 9月 19, 2022
  
  f4ec1563
16 9月, 2022 1 次提交

[Cherry-pick] Normalize yaml name and label (#46052) · 8caaf85a

由 Chen Weihang 提交于 9月 16, 2022

* normalize yaml file name (#45894)

* Clear extra attributes of activation op in OpMaker (#45772)

* clear extra attr of activation op in opmaker

* fix syntax bug

* fix mkldnn kernel

* fix merge conflict

* fix bug

* [PHI] Normalize yaml op label (#45976)

* normalize yaml op label

* revert op_compat yaml change

* fix prelu and rnn compat problem

* replace api by op

* support assign op backward refuse forward (#45879)

* normize yaml backward op label (#46028)
Co-authored-by: Nzyfncg <zhangyunfei07@baidu.com>
Co-authored-by: NCharles-hit <56987902+Charles-hit@users.noreply.github.com>

8caaf85a

15 9月, 2022 2 次提交
- W
  Support 0 shapes input Tensor for MKL slice (#45930) (#46072) · 903c87bd
  由 WangZhen 提交于 9月 15, 2022
```
Support 0 shapes input Tensor for MKL slice kernel
```
  903c87bd
- Z
  Delete eigen header in data_type.h (#46036) (#46066) · 2680a71e
  由 zyfncg 提交于 9月 15, 2022
```
* delete eigen header in data_type.h

* fix complie bug

* refactor
```
  2680a71e
14 9月, 2022 3 次提交
- J
  
  merge python lib (#46013) · 5130b0a1
  由 JingZhuangzhuang 提交于 9月 14, 2022
  
  5130b0a1
- L
  
  set device id before op run (#45994) · 2fac8abb
  由 Leo Chen 提交于 9月 14, 2022
  
  2fac8abb
- P
  
  delete new executor log (#45917) · e223cf7b
  由 pangyoki 提交于 9月 14, 2022
  
  e223cf7b
13 9月, 2022 2 次提交
- J
  
  cherry pick softmax infer kernel (#45957) · 0903020d
  由 JingZhuangzhuang 提交于 9月 13, 2022
  
  0903020d
- R
  [cherry-pick] Allow manaully set py_reader name in standalone executor (#45898) (#45931) · 29c44eb2
  由 Ruibiao Chen 提交于 9月 13, 2022
```
* Allow manaully set py_reader name in standalone executor

* Fix CI errors
```
  29c44eb2
09 9月, 2022 2 次提交
- R
  [CustomDevice] add dy2static support (#45878) · abc85c50
  由 ronnywang 提交于 9月 09, 2022
```
* [CustomDevice] add dy2static support

* update
```
  abc85c50
- C
  [Phi] Add fusion kernel dir and migrate fused_softmax_mask op (#45802) · 2b4f44d5
  由 Chen Weihang 提交于 9月 09, 2022
```
* add fusion dir and fuse_softmax_mask kernel

* remove fusion kernel dir

* migrate infershape

* fix code errror
```
  2b4f44d5
08 9月, 2022 3 次提交
- H
  
  polish code comment, test=doc (#45859) · 447d79da
  由 HongyuJia 提交于 9月 08, 2022
  
  447d79da
- A
  [OpAttr]Refine Teller logic if encounter OpDesc with Variable type Attribute (#45795) · a642365e
  由 Aurelius84 提交于 9月 08, 2022
```
* [OpAttr]Refine Teller logic if encounter OpDesc with Variable type Attribute

* fix iterator

* fix typo

* fix lambda expr

* fix ptr
```
  a642365e
- X
  [Dy2Static] Filter int64/int32/int16/bool in conditional op (#45759) · 36046a89
  由 xiongkun 提交于 9月 08, 2022
```
* stop pass filter int32/int16/int64/bool inputs in cond_op

* fix bugs: except block 0, the backward vars and forward vars exist in different blocks.

* fix code by review
```
  36046a89
07 9月, 2022 6 次提交
- L
  
  use xxhash instead of cryptopp (#45837) · a89e48fe
  由 Leo Chen 提交于 9月 07, 2022
  
  a89e48fe
- Y
  
  rename the template type name for tranpose (#45834) · 9b70c556
  由 Yuang Liu 提交于 9月 07, 2022
  
  9b70c556
- C
  [Phi] Fix infermeta bug for vector input and output (#45810) · 420d186a
  由 Chen Weihang 提交于 9月 07, 2022
```
* fix infermeta bug for vector input and output

* add unittest
```
  420d186a
- Y
  
  [alphafold] Transpose support large tensors where there numel is bigger than INT32_MAX (#45753) · d9a9e638
  由 Yuang Liu 提交于 9月 07, 2022
  
  d9a9e638
- W
  Layernorm shift partition (#45736) · 960109af
  由 wenbin 提交于 9月 07, 2022
```
* first commit

* conver done

* correct format

* layernorm_shift_partition

* correct convert

* redefine plugin

* runable

* bug fix

* modify ShiftPartitionPattern

* correct

* add UT

* modify ut

* compile

* modify enforce

* modify UT
```
  960109af
- C
  [Auto Parallel] Support Iterable dataset for auto parallel (#45518) · b77fa1d9
  由 caozhou 提交于 9月 07, 2022
```
* support iterable dataset for auto parallel

* add split_data proto

* fix unittest bug

* fix recompute bug

* update cmake
```
  b77fa1d9
06 9月, 2022 2 次提交
- Y
  [PHI]Add TensorArray for PHI (#45479) · 68f99b78
  由 YuanRisheng 提交于 9月 06, 2022
```
* add tensor array

* fix ci bugs

* fix ci bugs

* fix ci bugs

* fix ci bugs

* update by comment

* update code
```
  68f99b78
- D
  
  fix cmake download program (#45800) · 3f3f923b
  由 danleifeng 提交于 9月 06, 2022
  
  3f3f923b

BaiXuePrincess / Paddle 与 Fork 源项目一致

BaiXuePrincess / Paddle
与 Fork 源项目一致