- 02 2月, 2021 1 次提交
-
-
由 Shang Zhizhou 提交于
* add dla * add python api Co-authored-by: Nshangzhizhou <root@szth-rp-fanyi-opera49.szth.baidu.com> Co-authored-by: Nshangzhizhou <root@szth-rp-fanyi-opera49.szth.baidu.com>
-
- 27 1月, 2021 1 次提交
-
-
由 Wojciech Uss 提交于
Co-authored-by: NJacek Czaja <jacek.czaja@intel.com>
-
- 22 1月, 2021 1 次提交
-
-
由 Pei Yang 提交于
-
- 21 1月, 2021 1 次提交
-
-
由 QingshuChen 提交于
-
- 20 1月, 2021 3 次提交
-
-
由 AshburnLee 提交于
* Add tf32 support for A100 tensor core acceleration for cuBLAS (#28732) * Fixed an error * Fixed an error
-
由 AshburnLee 提交于
This PR is cherry-picked from PR: #29192 Function: Added TF32 switch for cuDNN. Turned on as default, turned off when users set the switch as False
-
由 Wilber 提交于
-
- 19 1月, 2021 11 次提交
-
-
由 pangyoki 提交于
Cherry pick PR #30520 . Fix error message of Inplace strategy.
-
由 Leo Chen 提交于
[cherry-pick] support layer_norm fp16 in dygraph amp (#30430)
-
由 Zhou Wei 提交于
cherry-pick #30553 fix bug of multicard grad ncclAllReduce, the gradient accumulater of parameters should be keep order, otherwsie, it will influence multicard ncclAllReduce of grad.
-
由 liym27 提交于
cherry-pick #30536
-
由 Zhen Wang 提交于
Fix the compiling error of update_loss_scaling when using cuda9.
-
由 hutuxian 提交于
-
由 hutuxian 提交于
-
由 tangwei12 提交于
* add trainers for pserver Change-Id: I99c0ab1cc427318f1f9bf8f8f5faff2b8890645d * add trainers for pserver Change-Id: I1a75793ec81ce126d07f4c47cae09b95d530bbc8
-
由 taixiurong 提交于
* support transformer v2.0 * fix range op crash in dygraph xpu place
-
由 liuyuhui 提交于
-
由 JZ-LIANG 提交于
-
- 18 1月, 2021 5 次提交
-
-
由 lidanqing 提交于
Co-authored-by: NWojciech Uss <wojciech.uss@intel.com>
-
由 guofei 提交于
* Modify the calculation logic of LambOptimizer (#29313) * Modify the calculation logic of LambOptimizer * Modify the calculation logic of LambOptimizer * Modify the calculation logic of LambOptimizer
-
由 ceci3 提交于
* add pad and concat double grad * resolve conflict
-
由 Zhang Ting 提交于
* add fp16 support for tril_triu op (#30186) * add VecCastCUDAKernel (#30296) Co-authored-by: Nfurnace <34057289+windstamp@users.noreply.github.com>
-
由 pangyoki 提交于
Cherry-pick PR 30103. Add Inplace strategy (Output reuse Input Varbase) in dygraph (#30103) (#30496) * add view strategy on squeeze,unsqueeze,reshape,flatten * add squeeze unittest * add unittests * use View strategy as name rather than Reuse Allacation * fix view api doc * fix format * use core.ops when input of reshape2 is Tensor * fix test_cross_entropy_loss error because of reshape2 * fix test_cross_entropy_loss error because of reshape2 * add inplace strategy * add elementwise_add sub * let backward op not use inplace * grad op do not use inplace * fix memory increase error and add leaf error message * delete selected_rows * change op_function * little change * solve HandleViewBetweenInputAndOutput * add unittest and leaf error message * merge view error * optimize op_function_generator format and support sum inplace op * fix format of basic_engine * fix format for framework * little change of variable wrapper * add reshape, squeeze, unsqueeze, scatter api * add relu elu tanh softmax inplace api * fix test_squeeze_op unittest * fix test_relu_op unittest * fix comment problems * delete sample code of inplace api * add reference of grad_pending_nodes in basic_engine * fix unittest name * add inplace apis into wlist * fix error message * add PADDLE_ENFORCE for set grad op twice * fix head file error
-
- 15 1月, 2021 6 次提交
-
-
由 pangyoki 提交于
* Cherry-pick 30072, add dispenable input for core.ops.reshape2/expand/slice (#30072) * add dispenable input 'shape' for core.ops.reshape2 * add dispenable inputs for core.ops.reshape2/expand/slice * add ut * save reshape update in pr 30180 * save reshape update v2 in pr 30180 Co-authored-by: NLeo Chen <chenqiuliang@baidu.com>
-
由 Yang Zhang 提交于
built-in `rsqrt` is shadowed
-
由 lijianshe02 提交于
* add transpose double grad test=develop (#29600) * add transpose double grad test=develop * cherry-pick test=develop
-
由 whs 提交于
-
由 123malin 提交于
* test=develop, add distributed_infer (#30300) * test=develop, add distributed_infer * test=develop, fix unittest cmakefile conflict * test=develop, fix test_dist_fleet_base
-
由 wawltor 提交于
* fix the rnn mask memory bug for out of read * update the code for the rnn
-
- 14 1月, 2021 8 次提交
-
-
由 ShenLiang 提交于
-
由 lidanqing 提交于
-
由 cc 提交于
-
由 QingshuChen 提交于
* optimize memcpy perf for kunlun (#30291) * optimize memcpy perf for kunlun * remove useless unitest for kunlun mean * minor * fix bug that cann't find mkldnn(kunlun) (#30394)
-
由 LielinJiang 提交于
* Add double grad for conv_transpose (#29706) * add double grad for conv_transpose * register cudnn conv double grad for depthwise conv (#29807)
-
由 Zhang Jun 提交于
-
由 alncat 提交于
-
由 GaoWei8 提交于
* softmax backward optimize
-
- 13 1月, 2021 3 次提交