1. 30 12月, 2021 9 次提交
    • Y
      [Auto parallel] Make sure the id semantics of every var and op unique (#38132) · 5620214e
      Yulong Ao 提交于
      * [Auto parallel] Make the id of var and op unique
      
      * [Auto Parallel] Rename back dist_context to distop_context
      5620214e
    • H
      Add cpu kernel of new api : lstsq (#38585) · ccf99b66
      Haohongxiang 提交于
      * add cpu kernel of lstsq
      
      * update
      
      * modify code style
      
      * modify unittest
      
      * remove support for complex
      ccf99b66
    • J
      Support test imperative basic with fixed retain grad interface (#38548) · 2421a25a
      Jiabin Yang 提交于
      * Rearranged Eager AutoCodeGen directory structure
      
      * Removed USE_OP in Eager AutoCodeGen
      
      * Enabled generation for Operators without Grad/Inputs/Outputs
      
      * Resolved operators without input
      
      * Fixed merge conflicts
      
      * Enabled Eager AutoCodeGen for 10+ more operators
      
      * Refactored Eager AutoCodeGen with more organized helper objects
      
      * Enabled Eager AutoCodeGen for operators with multiple OpBases
      
      * Adjusted Eager AutoCodeGen to Enable Passing Output Tensor as Input Argument
      
      * Handled Dispensable Inputs/Outputs in Eager AutoCodeGen
      
      * Adjusted function generation/call between Python-C API & Dygraph API
      
      * Synchronized auto-generated Python-C API with Dygraph Forward Functions
      
      * support more eager tensor api
      
      * fix merge compile error
      
      * fix compile error and fit develop code
      
      * support pure CPU
      
      * fix some logic error in eager_mode
      
      * support _varbase_creator in eager mode
      
      * Added safe_initialized interface to EagerTensor for use in processing dispensable inputs
      
      * for eager mode
      
      * refine
      
      * support multiple constructor for eager tensor
      
      * add place related code
      
      * polish code
      
      * specific randint with dtype of int64
      
      * Support pure cpu test
      
      * eager logic
      
      * refine test in pure cpu
      
      * eager logic
      
      * eager logic
      
      * eager logic, test=develop
      
      * skip core.eager when in inference, test=develop
      
      * refine, test=develop
      
      * refine, test=develop
      
      * call RetainGrad after run forward kernel, test=develop
      
      * refine, test=develop
      
      * support dygraph util, meta, guard test
      
      * support inference test
      
      * refine test and fix initializer failed
      
      * support create varbase and fix retain grad error
      
      * fix windows error
      
      * support test_imperative_basic test in eager mode
      
      * remove additional log in variable.h
      
      * remove additional log in variable.h
      
      * remove additional code create in merge
      Co-authored-by: Njim19930609 <jim19930609@gmail.com>
      Co-authored-by: NWang Huan <wanghuan29@baidu.com>
      2421a25a
    • J
      Added Conv2D BF16 BWD oneDNN kernel (#38507) · ed8ba011
      jakpiase 提交于
      * working test for padding only
      
      * added full conv2d grad kernel
      
      * removed some trash
      
      * minor change
      
      * Ci fix
      
      * format fix
      ed8ba011
    • Z
      [PSCore]Fix test fleet base 2 (#38588) · 04496d89
      zmxdream 提交于
      04496d89
    • C
      [PTen] Remove offset in storage (#38472) · a504ff3f
      Chen Weihang 提交于
      * remove offset in storage
      
      * revert api change
      
      * fix custom op slice bug
      
      * fix mutable_data error
      a504ff3f
    • X
      add ExponentialFamily and Dirichlet probability distribution (#38445) · 00cddf07
      Xiaoxu Chen 提交于
      * extend Distribution baseclass for supporting multivariant distribution and prob method
      
      * add ExponentialFamily base class and entropy using Bregman divergence
      
      * add dirichlet probability distribution
      00cddf07
    • X
      add dirichlet random sample op in cpu and gpu kernel (#38244) · c5bf09bb
      Xiaoxu Chen 提交于
      * add dirichlet sample op and cpu backend kernel
      
      * add Dirichlet op cuda kernel  (#6)
      
      * add dirichlet op hip kernel
      Co-authored-by: NFeiyu Chan <chenfeiyu@baidu.com>
      c5bf09bb
    • L
      Fix the bug of batch_norm and batch_norm_grad op. (#38288) · cc83c95f
      Leo Guo 提交于
      * Fix the bug of batch_norm and batch_norm_grad op. Add the "roi_align" and "roi_align_grad" op in xpu2 op list.
      
      * Fix the bug of batch_norm and batch_norm_grad op. Add the "roi_align" and "roi_align_grad" op in xpu2 op list. test=kunlun
      Co-authored-by: NZibin <guozibin@baidu.com>
      cc83c95f
  2. 29 12月, 2021 11 次提交
  3. 28 12月, 2021 15 次提交
    • Z
      add new API: paddle.cov (#38392) · 85f5d264
      zhiboniu 提交于
      85f5d264
    • B
      update seq_concat_fc_fuse_pass ut (#38538) · 706d2c08
      baoachun 提交于
      706d2c08
    • F
      Utilize StreamSafeCUDAAllocator to support fast GC in new executor (#37642) · 0c7153a4
      From00 提交于
      * fix reshape move storage error
      
      * remove needless set type
      
      * alloc tensor by shared storage
      
      * Utilize StreamSafeCUDAAllocator to support fast GC in new executor
      
      * Fix compile error for Windows and ROCm
      
      * Fix compile error for Windows
      
      * Modify UT stream_safe_cuda_alloc_test
      
      * Modify UT stream_safe_cuda_alloc_test
      
      * Rewrite fast GC
      
      * Rewrite fast GC
      
      * Fix compile error for BOOST_GET_CONST
      
      * Fix compile error for BOOST_GET_CONST
      
      * Changes default stream for StreamSafeCUDAAllocator
      
      * Fix a small CI error
      
      * Remove some redundant code
      
      * Fix conflict
      
      * Fix compile error for ROCm
      
      * Fix Windoes CI error
      
      * Fix CI error
      
      * Remove some unnecessary code
      
      * Fix CI error
      
      * Add UT for fast GC
      
      * Fix CI error
      
      * add device-agnostic stream class
      
      * add stream.h
      
      * fix ut
      
      * fix cpu compile
      
      * Use RWLock in GetAllocator
      
      * Fix CI error
      Co-authored-by: NChen Weihang <chenweihang@baidu.com>
      Co-authored-by: Nzhiqiu <chenqiuliang@baidu.com>
      0c7153a4
    • H
      add matmul_to_mul matmul_v2_to_mul matmul_v2_to_matmul test case (#37645) · bed71992
      heliqi 提交于
      * add matmul_to_mul matmul_v2_to_mul matmul_v2_to_matmul test case
      
      * modify skip func to ignore_pass_case func
      
      * rebuild CI
      
      * rebuild CI
      
      * add test_map_xx_pass timeout
      
      * add test_map_xx_pass timeout
      
      * merge from develop
      
      * add timeout notest;test=coverage
      
      * Cmakelist add timeout
      
      * add timeout
      
      * add attr of matmul_v2
      
      * add trt skip
      
      * delete trt config
      
      * add skip,  mul diff on 3080
      bed71992
    • T
      refine amax/amin example(#38525) · 00a50af8
      Tao Luo 提交于
      00a50af8
    • J
      Support test basic of Var and Layer (#38426) · 1fb80a6a
      Jiabin Yang 提交于
      * Rearranged Eager AutoCodeGen directory structure
      
      * Removed USE_OP in Eager AutoCodeGen
      
      * Enabled generation for Operators without Grad/Inputs/Outputs
      
      * Resolved operators without input
      
      * Fixed merge conflicts
      
      * Enabled Eager AutoCodeGen for 10+ more operators
      
      * Refactored Eager AutoCodeGen with more organized helper objects
      
      * Enabled Eager AutoCodeGen for operators with multiple OpBases
      
      * Adjusted Eager AutoCodeGen to Enable Passing Output Tensor as Input Argument
      
      * Handled Dispensable Inputs/Outputs in Eager AutoCodeGen
      
      * Adjusted function generation/call between Python-C API & Dygraph API
      
      * Synchronized auto-generated Python-C API with Dygraph Forward Functions
      
      * support more eager tensor api
      
      * fix merge compile error
      
      * fix compile error and fit develop code
      
      * support pure CPU
      
      * fix some logic error in eager_mode
      
      * support _varbase_creator in eager mode
      
      * Added safe_initialized interface to EagerTensor for use in processing dispensable inputs
      
      * for eager mode
      
      * refine
      
      * support multiple constructor for eager tensor
      
      * add place related code
      
      * polish code
      
      * specific randint with dtype of int64
      
      * Support pure cpu test
      
      * eager logic
      
      * refine test in pure cpu
      
      * eager logic
      
      * eager logic
      
      * eager logic, test=develop
      
      * skip core.eager when in inference, test=develop
      
      * refine, test=develop
      
      * refine, test=develop
      
      * call RetainGrad after run forward kernel, test=develop
      
      * refine, test=develop
      
      * support dygraph util, meta, guard test
      
      * support inference test
      
      * refine test and fix initializer failed
      
      * support create varbase and fix retain grad error
      
      * fix windows error
      
      * support test code coverage
      
      * support test code coverage
      
      * support test code coverage
      Co-authored-by: Njim19930609 <jim19930609@gmail.com>
      Co-authored-by: NWang Huan <wanghuan29@baidu.com>
      1fb80a6a
    • W
      fix ci problem (#38474) · 2e4cb279
      Wilber 提交于
      2e4cb279
    • H
      Add API and op for take_along_axis (#38396) · 3310f519
      huangxu96 提交于
      * add API and op for take_along_axis
      
      * fix compile dependency problem and add example code and doc
      
      * add unitest
      
      * delete some code for CI coverage
      
      * fix code style problem
      
      * fix as review
      3310f519
    • T
      Add Amax and Amin API (#38417) · 340dfb26
      Tao Luo 提交于
      * add amax/amin
      
      * support axis is list
      340dfb26
    • C
      [pten] remove in_type arg in cast kernel (#38486) · 0637b9a6
      chentianyu03 提交于
      * remove intype arg in cast kernel
      
      * modify conj config in api.yaml by dictionary order
      
      * rm unused code in cast_kernel.cu
      0637b9a6
    • H
      add reduce_prod_xpu. fix reduce_mean_xpu bug. (#38481) · 78836bb7
      houj04 提交于
      * add reduce_prod_xpu. fix reduce_mean_xpu bug.
      
      * iadd reduce_prod_xpu. fix reduce_mean_xpu bug. test=kunlun
      78836bb7
    • B
      add mul_lstm_fuse_pass ut (#37795) · 1db61c3e
      baoachun 提交于
      * add mul_lstm_fuse_pass ut
      
      * update mul_lstm_fuse_pass ut
      
      * update ut
      
      * update ut
      
      * update ut
      
      * add CPU ut cmake setting
      
      * update ut
      1db61c3e
    • Z
      add pass base unittest (#38504) · ee5f3641
      zhaoyingli 提交于
      * add pass base unittest
      
      * update gpt model
      ee5f3641
    • S
      fix compile dir conflict with include_dirs (#38479) · e42ed7d1
      sneaxiy 提交于
      e42ed7d1
    • L
      Fix scatter_op fp16 perf problem. (#38499) · 33ce249f
      Li Min 提交于
      * Fix scatter_op fp16 perf problem.
      
      * Add scatter into black list.
      
      * Add scatter into black list for dygraph.
      33ce249f
  4. 27 12月, 2021 5 次提交