提交 · 10f0a0f6c8f71436bad715b0f74329e89ea076f9 · 机器未来 / Paddle

18 10月, 2021 13 次提交

[HybridParallel]Support fp16 in dygraph hybrid parallel (#36420) · 10f0a0f6

由 Haohongxiang 提交于 10月 18, 2021

* [HybridParallel]Support fp16 in dygraph hybrid parallel

* update

* update

* update for recompute

* add unittest of pp+fp16

* add unittest of recompute+fp16

* update

* modify ut

10f0a0f6

Added softplus FP32 FWD OneDNN kernel (#36382) · bdac9ff6

由 jakpiase 提交于 10月 18, 2021

* added softplus

* refactored softplus op

* deleted unnecessary file

* added missing file

* added formatting

* disabled tests if GPU is used

* added reviewer suggestion

* unified softplus kernel

bdac9ff6

Lml/vhp (#36146) · 4c0ad772

由 levi131 提交于 10月 18, 2021

* init functional jacobian api

* finish test with dtype float32

* add float64 test case

* polish code

* use atol=1e-5 with dtype float64

* fix for ci

* set timeout for test_jacobian

* init hessian API

* save status

* polish API docstring

* modify docstring

* add utils.py

* save status

* fix dygraph double grad dtype error when calling for high differential senario

* reinvoke ci

* test_hessian.py is ok

* polish hessian API

* init vhp

* Revert "init vhp"

This reverts commit cbd4d3b66abe82b0ac10721b9eddeb7d82e0a1c8.

* init vhp

* finish vhp API logically

* add test for partial_engine.cc

* modify numerical_delta with dtype float32

* merge fix for dtype float64

* spell fix

* save status

* polish code

* rm _stop_gradient_pre_process

* save status

* add example for vhp interface

* add _compute_numerical_vjp and _compute_numerical_vhp

* test is ok

* vhp is ok

* add testVHPFloat64

* modify for comments

* modify format

* modify format

* save status

* test_vhp is ok

* finish code polish

* small modify for v is None
Co-authored-by: NJiabinYang <360788950@qq.com>

4c0ad772

Add quant axis (#36467) · b7f76647

由 xiaoxiaohehe001 提交于 10月 18, 2021

* add_quant_axis

* add_quant_axis

* --amend

* Update quant_conv2d_dequant_fuse_pass.cc

b7f76647

Q

[NPU] add kernels for elementwise_add gather_nd tile, test=develop (#36464) · cbd15f7d
由 Qi Li 提交于 10月 18, 2021

cbd15f7d
Q

[NPU] fix dtype for arg_max, test=develop (#36457) · 8757fc5b
由 Qi Li 提交于 10月 18, 2021

8757fc5b

Add operators for async read & async write (#36333) · 3845afff

由 Siming Dai 提交于 10月 18, 2021

* fix async_read bug

* change index place to cpu

* add tensor size judge

* add async_read & async_write test

* fix bug in async_write

* fix mac py3 ci

* fix bug for cpu version paddle

* fix windows ci bug

* change input argument error type

* change const_cast to mutable_data

* add async_write out-of-bound check and consumate error hint

* fix a small bug for dst_tensor

* add docs and refine codes

* refine docs

* notest,test=windows_ci

* fix windows ci

* fix require

* fix code-block

* add core.is_compiled_with_cuda()

3845afff

C
quant support matmul_v2 (#36469) · 051544b6
由 ceci3 提交于 10月 18, 2021
```
* quant support matmul_v2

* fix format
```
051544b6
W

add IPluginV2Layer: AddPluginV2Ext (#36493) · 623e36b0
由 Wangzheee 提交于 10月 18, 2021

623e36b0
T
[XPU AMP] 1. xpu support gradient acc 2. xpu support create tensor in dygraph... · d19a9b39
由 taixiurong 提交于 10月 18, 2021
```
[XPU AMP] 1. xpu support gradient acc 2. xpu support create tensor in dygraph 3. xpu support update weight params in amp (#36439)
```
d19a9b39
J

Fix conv2d op_teller error (#36474) · d3c93942
由 JingZhuangzhuang 提交于 10月 17, 2021

d3c93942

[autograd.functional] Fix a bug on handling v=None in vjp and jvp (#36445) · 79dbbcce

由 Tongxin Bai 提交于 10月 18, 2021

* autograd.functional passed pylint checker.

* autograd.functional: fix import errors.

* autograd.functional: fixed unit tests.

* autograd.functional minor format change

* [autograd.functional] Fixed vjp and jvp's v=None bug.

79dbbcce

H

modify ut of cond (#36475) · e496d1e9
由 Haohongxiang 提交于 10月 18, 2021

e496d1e9

17 10月, 2021 2 次提交
- Z
  
  refine rescale_grad (#36490) · 4e036fa1
  由 Zeng Jinle 提交于 10月 17, 2021
  
  4e036fa1
- Z
  Revert "fix the initializer of resnet unit op (#36483)" (#36487) · 314cc495
  由 Zeng Jinle 提交于 10月 17, 2021
```
This reverts commit 0452f27c.
```
  314cc495
16 10月, 2021 1 次提交
- Z
  fix the initializer of resnet unit op (#36483) · 0452f27c
  由 Zhang Zheng 提交于 10月 16, 2021
```
* fix the initializer of resnet unit op

* fix the initializer of resnet unit op
```
  0452f27c
15 10月, 2021 11 次提交

Z
Remove wrong __restrict__ of CUDA LarsMomentumOpKernel (#36460) · adb80494
由 Zeng Jinle 提交于 10月 15, 2021
```
* remove wrong restrict

* remove master_param_out __restrict__

* update
```
adb80494
D

fix opt-offload save bug (#36433) · e703a2ed
由 duanboqiang 提交于 10月 15, 2021

e703a2ed
Z

Add ResNetUnit Python API (#35426) · 12882b2f
由 Zhang Zheng 提交于 10月 15, 2021

12882b2f
F

feat: Add TRT support for 3D(batch_norm_op and elementwise_add_op) (#36446) · 2de0b58e
由 feng_shuai 提交于 10月 15, 2021

2de0b58e

add resnext (#36070) · 277c9a55

由 Nyakku Shigure 提交于 10月 15, 2021

* add resnext model
* add zh docs
* add unittest
* test performance
Co-authored-by: Ainavo <ainavo@163.com>
Co-authored-by: Npithygit <pyg20200403@163.com>
Co-authored-by: Ainavo <ainavo@163.com>
Co-authored-by: Npithygit <pyg20200403@163.com>

277c9a55

0
fix no_grad context error in train mode when using save/load (#36434) · 37257d6a
由 0x45f 提交于 10月 15, 2021
```
* fix no_grad context error in train mode when using save/load

* change net to train mode in test case
```
37257d6a
F

dynamic load mkl as a fft backend when it is avaialble and requested (#36414) · f45e6cf6
由 Feiyu Chan 提交于 10月 15, 2021

f45e6cf6

Add BuildCinnPass (#36345) · b3f02c57

由 jiangcheng 提交于 10月 15, 2021

* Add CinnSubgraphSearchPass

* solve CI problem of subgraph order not same

* fix some bug by review advices

* ensure the independently of subgraph, that mean the subgraph should not have link to out-graph

* rename cinn_subgraph_search_pass to build_cinn_pass and delete paddle_to_cinn_pass

* add flag to control wheter append build cinn pass

* remove AppendPass at ParallelExecutorPassBuilder

* rename paddle_to_cinn_pass to build_cinn_pass in build_strategy and close test_run_from_cinn

b3f02c57

[New Feature] Support tanh triple grad (#36225) · 808be657

由 Jiabin Yang 提交于 10月 15, 2021

* native commit for triple grad of sigmod

* Updated unittests files

* init functional jacobian api

* Updated trible_test func

* Updated gradient_checker & test_script

* finish test with dtype float32

* add float64 test case

* polish code

* use atol=1e-5 with dtype float64

* fix for ci

* set timeout for test_jacobian

* fix dygraph grad to support high differential

* polish API docstring

* Updated gradient checker and some related files

* fix double grad strip error for high differential

* fix double grad strip error for high differential

* Add Sigmoid triple grad tests

* fix dygraph double grad dtype error when calling for high differential senario

* Updated triple grad teses func

* Use np.random to initialize ddx

* Updated triple_grad_check func

* add todo for gradient checker and refine some comments

* remove additional code

* add test for warnging in backward.py

* add tanh triple grad

* format python code

* refine code
Co-authored-by: Nveyron95 <veyron_wu@163.com>
Co-authored-by: Nlevi131 <limaolin01@baidu.com>

808be657

Z

fix momentum ops (#36452) · 4dda18a8
由 Zeng Jinle 提交于 10月 15, 2021

4dda18a8
W

close some check on CI-OP-Benchmark, test=develop (#36442) · 8566cc98
由 wuhuanzhou 提交于 10月 15, 2021

8566cc98

14 10月, 2021 13 次提交
- Y
  add sparse_embedding doc (#36283) · 6ccc2a40
  由 Yanxing Shi 提交于 10月 14, 2021
```
* add sparse_embedding doc

* delete wrong space

* fix error for sample code

* fix error for doc compile

* delete __all__

* modify sample code
```
  6ccc2a40
- D
  
  optimize-offload support adamw op type (#36432) · 66c58fa3
  由 duanboqiang 提交于 10月 14, 2021
  
  66c58fa3
- Z
  
  fix lars (#36431) · 8256f6fa
  由 Zeng Jinle 提交于 10月 14, 2021
  
  8256f6fa
- L
  
  enable 3rd order test case (#36427) · 3cf57646
  由 levi131 提交于 10月 14, 2021
  
  3cf57646
- Z
  
  refine merge lars (#36428) · 63fd7d66
  由 Zeng Jinle 提交于 10月 14, 2021
  
  63fd7d66
- W
  inference support bert when exists matmul_v2 (#36424) · 3e6d9dbb
  由 Wilber 提交于 10月 14, 2021
```
* support bert when exists matmul_v2

* update
```
  3e6d9dbb
- Z
  
  Add the complete code and related files of resnet_unit_op (#36366) · 12e6dbbc
  由 Zhang Zheng 提交于 10月 14, 2021
  
  12e6dbbc
- Z
  [NPU] Add density_prior_box (#36361) · bed4fb27
  由 zhulei 提交于 10月 14, 2021
```
* [NPU] Add density_prior_box op

* [NPU] Add density_prior_box op
```
  bed4fb27
- L
  Revert "Implemented LRU based cache clearing (#36290)" (#36426) · 5d18967b
  由 lidanqing 提交于 10月 14, 2021
```
This reverts commit bf748f24.
```
  5d18967b
- Z
  Merge momentum ops/kernels (#36380) · f4eda869
  由 Zeng Jinle 提交于 10月 14, 2021
```
* merge momentum ops

* update

* add ut to improve coverage

* remove optimizer change

* fix error msg

* update ut

* add __restrict__ for CUDA

* update ut

* move merged_momentum_op to optimizer dir

* fix coverage
```
  f4eda869
- Z
  
  refine lars (#36409) · eb722e34
  由 Zeng Jinle 提交于 10月 14, 2021
  
  eb722e34
- S
  [HybridParallel]Rebuild code for pipeline (#36396) · 8ffcc7c8
  由 ShenLiang 提交于 10月 14, 2021
```
* add no_sync for parameters sync

* add pipeline for moe
```
  8ffcc7c8
- S
  
  reduce some unittest's parallel number to avoding timeout failure (#36397) · 693b1aa1
  由 Sing_chan 提交于 10月 14, 2021
  
  693b1aa1

机器未来 / Paddle 与 Fork 源项目一致

机器未来 / Paddle
与 Fork 源项目一致