提交 · 3a81805bbce4ea4a01a81d06c626c20fa9cfeed9 · 机器未来 / Paddle

26 11月, 2021 2 次提交

add new API/OP: paddle.linalg.triangular_solve (#36714) (#37551) · 3a81805b
由 zhouweiwei2014 提交于 11月 26, 2021
```
cherry-pick #36714
```
3a81805b

[cherry-pick 2.2 heterps]bug fix for launch_utils.py (#37521) (#37570) · 4b41b8e9

由 zmx 提交于 11月 26, 2021

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix. test=develop

* [heterps]bug fix for _run_from_dataset

* fix heter_server.cc

* fix launch_utils.py

* fix heter_section_worker.cc

* fix. test=develop

* fix. test=develop

4b41b8e9

25 11月, 2021 3 次提交
- S
  [cherry-pick 2.2]fix data parallel when VOCAB var in program (#37546) · c8429d36
  由 Steffy-zxf 提交于 11月 25, 2021
```
* fix data parallel when VOCAB var in program

* fix ci coverage
```
  c8429d36
- K
  
  avoid setting logging.basicConfig (#37031) (#37530) · 824c4ef9
  由 kuizhiqing 提交于 11月 25, 2021
  
  824c4ef9
- P
  Cherry-pick PR 37420, fix inplace bug when the first grad_var(loss_grad) is... · d31d597f
  由 pangyoki 提交于 11月 25, 2021
```
Cherry-pick PR 37420, fix inplace bug when the first grad_var(loss_grad) is inplace var (#37420) (#37488)

fix inplace bug，Cherry pick PR #37420
```
  d31d597f
24 11月, 2021 1 次提交
- L
  [Cherry pick 2.2] fix bugs to support bias add none for fused_attention op. (#37411) (#37483) · bed652d6
  由 Li Min 提交于 11月 24, 2021
```
Add support for bias is none for fused_attention op.
```
  bed652d6
23 11月, 2021 6 次提交

L

bug fix shard_index (#37042) (#37421) · f873d3a1
由 lilong12 提交于 11月 23, 2021

f873d3a1

[cherry-pick]Refactor Heterogenous Pipeline Parameter Server (#37446) · 4dc426f4

由 zmx 提交于 11月 23, 2021

* bug fix for  DeserializeSelectedRows. test=develop (#36520)

* fix SerializeSelectedRows (#36543)

* bug fix for  DeserializeSelectedRows. test=develop

* fix bug for SerializeSelectedRows. test=develop

* update. test=develop

* [Heterps]Refactor Heter Pipeline Parameter Server (#36845)

* change username

* fix

* fix

* fix

* fix

* fix

* update

* update

* update unittests

* fix

* update

* fix

* update

* fix

* fix

* fix

* update

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update send_and_recv op. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* update. test=develop

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix ut. test=develop

* fix unit. notest,test=coverage

* fix ut. notest, test=coverage

* update. notest,test=coverage

* fix ut. notest, test=coverage

* fix ut. notest, test=coverage

* fix. notest, test=coverage

* fix. notest, test=coverage

* fix ut. notest, test=coverage

* fix ut. notest, test=coverage

* fix ut. notest, test=coverage

* fix ut. notest, test=coverage

* add func. notest, test=coverage

* fix ut. notest, test=coverage

* fix. test=develop

* fix. test=develop

* Fix unit test for send_and_recv_cpu & send_and_recv_gpu (#37129)

* [heterps]fix ut for heter_pipeline_trainer.cc  (#37136)

* fix ut. test=develop

* fix ut. test=develop

* [heterps]bug fix for local training with --heter_worker_num (#37166)

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix ut. test=develop

* [heterps]Refactor heterogenous worker (#37244)

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix ut. test=develop

* refactor heter trainer. test=develop

* fix. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix ut. test=develop

* [heterps]add heterps mode judgement (#37298)

* [heterps]change default executor for heter trainer (#37314)

* fix pslib. test=develop

* add device to train_from_dataset. test=develop

* refine fleet.stop_worker. test=develop

* fix ut. test=develop

* fix ut. test=develop

* fix executor & ut. test=develop

* fix executor & ut. test=develop

* fix executor & ut. test=develop

* [heterps]remove api for heter pipeline ps (#37396)

* fix api. test=develop

* fix api. test=develop

* fix code style. test=release/2.2

* fix CMakeLists. test=develop (#37454)

4dc426f4

Z

elu support alpha < 0 (#37316) (#37437) · 436808c6
由 zhupengyang 提交于 11月 23, 2021

436808c6

cherry pick save/load in the_one_ps (#37461) · 58a51130

由 wangguanqun 提交于 11月 23, 2021

* save/load in ps runtime(the_one_ps) (#36097)

* add trainer desc config to distributed strategy

* code style modified

* data_feed set lod

* fix bug

* code style

* fix bug

* save load

* save load

* save unittest

* add unittest of the_one_ps

* unittest

* add todo in communicator sendsparse

* fix bug in save_inference_model (#37362)

58a51130

[Dy2stat]Allow users to switch eval/train mode when using @to_static to... · eed736dc

由 0x45f 提交于 11月 23, 2021

[Dy2stat]Allow users to switch eval/train mode when using @to_static to decorate a function (#37383) (#37432)

本PR之前使用@to_static装饰一个单独的function时，对于生成的Program无法切换train/eval模式，只能运行在train模式下。这也就导致动转静后用户多次调用function显存会一直增长。
本PR之后，使用@to_static装饰一个单独的function时，可以通过function.train()或者function.eval()的方式来切换train/eval模式。

eed736dc

W

fix shape api (#37412) · 2778fcd9
由 Wilber 提交于 11月 23, 2021

2778fcd9

22 11月, 2021 2 次提交
- C
  Fix a bug of quantization (#36982) (#37381) · 9ffb43be
  由 ceci3 提交于 11月 22, 2021
```
* fix a quantization bug
Co-authored-by: NXGZhang <46363693+XGZhang11@users.noreply.github.com>
```
  9ffb43be
- S
  [cherry-pick] Add paddle.incubate.graph_send_recv API(#37205) (#37343) · 109f8a8e
  由 Siming Dai 提交于 11月 22, 2021
```
* Add paddle.incubate.graph_send_recv API

* fix bug in CudaAtomicMin and CudaAtomicMax

* add empty line
```
  109f8a8e
19 11月, 2021 1 次提交
- 0
  [Dy2stat]Support `for i in [1,2,3]` statements in dy2stat (#37259) (#37356) · 44db219a
  由 0x45f 提交于 11月 19, 2021
```
该PR使得动转静模块能够正确转换如下的for i in [1, 2, 3]语句。
```
  44db219a
16 11月, 2021 2 次提交

[cherry-pick-2.2.1]fix fused_transformer_encoder_layer bug (#37229) · 36dd295e

由 zhangkaihuo 提交于 11月 16, 2021

修复了fused_transformer_encoder_layer fine-tune过程发现的一些问题：

    fused_attention_op添加attn_mask=None的支持：PR
    pre_layer_norm处理问题：PR
    参数处理，计算错误的问题：PR
    add_bias计算错误问题：PR
    添加pure fp16的支持：PR

36dd295e

Z
fix bug of indexing with ellipsis (#37192) · 79b9f47e
由 zyfncg 提交于 11月 16, 2021
```
修复了一维Tensor在使用省略号(...)索引时维度检测异常的问题。
```
79b9f47e

15 11月, 2021 1 次提交
- Z
  MLPerf Optimization for Release/2.2 (#37109) · 287ca7d5
  由 Zeng Jinle 提交于 11月 15, 2021
```
* add mlperf optimization PRs

* update
```
  287ca7d5
10 11月, 2021 1 次提交
- J
  Fix rnn grad bug in cpu when dropout is zero (#37080) (#37086) · 70cb0a54
  由 Jack Zhou 提交于 11月 10, 2021
```
* fix rnn grad bug when num_layers is set 2 and dropout_prob is set 0

* add more test for rnn
```
  70cb0a54
08 11月, 2021 2 次提交
- W
  Optimized the solve op code:renamed var and removed template func (#36981) (#37011) · a787b278
  由 Weilong Wu 提交于 11月 08, 2021
```
    Renamed the variable and function
    Removed the original template function
    Removed the tests_properties in CMakeLists.txt
```
  a787b278
- Z
  setitem support passing stop_gradient from value to tensor (#37028) · 76cab751
  由 zyfncg 提交于 11月 08, 2021
```
att,Fix issue:36902
```
  76cab751
01 11月, 2021 1 次提交
- L
  [cherry-pick]fix cusparse compile bug in CUDA11.2, test=release/2.2 (#36913) · ab2004bb
  由 Liu-xiandong 提交于 11月 01, 2021
```
* fix cusparse compile bug in CUDA11.2, test=develop

* fix bug
```
  ab2004bb
30 10月, 2021 1 次提交
- Y
  Move the ASP training API to paddle.static.sparsity. (#36525) (#36860) · 09bc9c06
  由 Yiqun Liu 提交于 10月 30, 2021
```
Cherry-pick #36525
```
  09bc9c06
29 10月, 2021 1 次提交
- F
  1. fix ifftshift(missing negative sign before shifts); (#36835) · fa7aa6b8
  由 Feiyu Chan 提交于 10月 29, 2021
```
2. add complex data type support for paddle.shape at graph assembly.
```
  fa7aa6b8
28 10月, 2021 10 次提交

0

polish _remove_no_value_return_var() function (#36826) (#36830) · c716cf35
由 0x45f 提交于 10月 28, 2021

c716cf35
P
【Cherry-pick PR 36511】fix out_of_range bug of multinomial op's cuda kernel (#36511) (#36808) · d8ffb261
由 pangyoki 提交于 10月 28, 2021
```
Cherry-pick PR #36511
```
d8ffb261
Z

fix dygraph adamw (#36745) (#36794) · e3db65d5
由 zhaoyingli 提交于 10月 28, 2021

e3db65d5
L
fix device docs;test=document_fix (#36784) (#36827) · 0b7f43ec
由 Ligoml 提交于 10月 28, 2021
```
* fix device docs;test=document_fix

* update __init__.py
```
0b7f43ec

Cherry-pick-36556: add paddle.version.cuda and paddle.version.cudnn API (#36556) (#36795) · 05b8630f

由 pangyoki 提交于 10月 28, 2021

* add paddle.version.cuda and paddle.version.cudnn API

* fix little bug

* fix bug

* add doc string

* fix mkdir error

* fix windows path

* fix new paddle/version path

* fix unittest

* fix format

05b8630f

[cherry-pick 2.2]support quantization of bert (#36820) · f20c5c9c

由 XGZhang 提交于 10月 28, 2021

* [cherry-pick 2.2]support quantization of bert

support quantization for maumul_v2

* Update quantization_pass.py

f20c5c9c

[Cherry-pick] Enable CTC grad compute on GPU (#36780) · 8ede9e6f

由 Hui Zhang 提交于 10月 28, 2021

* Revert "Align CTC grad scale same with ESPNet (#34729)"

This reverts commit 10f9644c.

* ctc grad compute on gpu

8ede9e6f

L
Fix fused_attention_op and fused_feedforward_op bug when pre_layer_norm is false. (#36793) (#36816) · ae592233
由 Li Min 提交于 10月 28, 2021
```
* Fix bug when pre_layer_norm is false.
```
ae592233

[Cherry-pick]FFT function enhancements and bugfixes (#36537) · 11b9f5f9

由 Xiaoxu Chen 提交于 10月 28, 2021

* update fft api path (#36219)

* update fft api path
* add sample code for ihfft2
Co-authored-by: Nchenfeiyu <chenfeiyu@baidu.com>

* fix fft axis (#36321)

fix: `-1` is used when fft's axis is `0`

* use unified external error message for cufft api (#36114)

* fft: modify sample code result (#36325)

* dynamic load mkl as a fft backend when it is avaialble and requested (#36414)

* add rocm support for fft api (#36415)

* move signal apis

* move fft and signal API path (#2)

* move signal apis

* move fft.py and signal.py to paddle/, fix typos

* fix relative imports from fft.py and signal.py

* fix typos in signal.py (#3)

* move signal apis

* move fft.py and signal.py to paddle/, fix typos

* fix relative imports from fft.py and signal.py

* fix typos

* disable Cache when CUFFT_VERSION >= 10200 (#4)

* move signal apis

* move fft.py and signal.py to paddle/, fix typos

* fix relative imports from fft.py and signal.py

* fix typos

* Add LRUCache for fft plans

* add LRUCache for cuff and hipfft (#5)

* move signal apis

* move fft.py and signal.py to paddle/, fix typos

* fix relative imports from fft.py and signal.py

* fix typos

* WIP: add cache

* delete move constructor and operator= for CuFFTHandle and FFTConfig

* remove log from CuFFTHandle and FFTConfig

* add lrucache for fft rocm backend

* disable LRUCache when CUFFT_VERSION >= 10200

* disbale copy and move for hipFFTHandle; format code
Co-authored-by: NXiaoxu Chen <chenxx_id@163.com>

* remove debug message of cufftHandler

* roll_op: support Tensor as input for shifts (#36727)

* fix fftshift/ifftshift on static mode

* update roll_op version

* add more test cases for fftshift/ifftshift
Co-authored-by: Nzhiboniu <31800336+zhiboniu@users.noreply.github.com>
Co-authored-by: Nchenfeiyu <chenfeiyu@baidu.com>
Co-authored-by: LJQ❤️ <33169170+lijiaqi0612@users.noreply.github.com>

11b9f5f9

0
show paddle traceback after last user code traceback (#36741) (#36765) · 96edcea4
由 0x45f 提交于 10月 28, 2021
```
show paddle traceback after last user code traceback
```
96edcea4

27 10月, 2021 4 次提交

Z
[cherry-pick]Fused transformer encoder layer and fused feedforward layer #36776 · e1b5b1da
由 zhangkaihuo 提交于 10月 27, 2021
```
本PR是fused_transformer的layer层代码，包含FusedFeedForward的layer层代码和FusedTransformerEncoderLayer的代码。
```
e1b5b1da
H
Modify paddle.static.nn.cond doc (#36694) (#36767) · c542d571
由 Huihuang Zheng 提交于 10月 27, 2021
```
Update `cond` English document
```
c542d571
H

cherrypick for eigvalsh (#36680) · 9d2e0923
由 huangjun12 提交于 10月 27, 2021

9d2e0923

Add fused attention op backward and python layer. (#36498) (#36752) · 64643d50

由 Li Min 提交于 10月 27, 2021

功能：本PR的目标是提高attention模块的计算性能。
为了减少框架层对op的调度开销，本PR通过在C++层手动实现attention模块，对外提供attention 大op；
为了减少防存开销，本PR采取了两种优化方法：
（1）在q,k,v计算时通过共享输入X，将该处的gemm，transpose和bias add从三次调用减少为一次；
（2）使用kernel融合优化技术，在不同cuda kernel之间通过寄存器传输数据；

64643d50

26 10月, 2021 2 次提交
- W
  
  [cherry-pick] enable trt test check and fix trt ut error（3/3） (#36696) · 211cf208
  由 Wilber 提交于 10月 26, 2021
  
  211cf208
- B
  fix wrong trt dim when input dim is 2 (#36614) (#36732) · da6e5143
  由 baoachun 提交于 10月 26, 2021
```
* fix wrong trt dim when input dim is 2

* update leaky_relu and instance_norm converter unit test

* add instance_norm input dim check
```
  da6e5143

机器未来 / Paddle 与 Fork 源项目一致

机器未来 / Paddle
与 Fork 源项目一致