提交 · b9fdd3bc0f4f22af17a81bb8a50a337b563c876b · PaddlePaddle / Paddle

01 11月, 2021 1 次提交

Paddle Tensor Operation Library initial implementation (#34425) · b9fdd3bc

由 Chen Weihang 提交于 11月 01, 2021

* initial tensor design & sign kernel demo

* add move constructor for meta & add lodtensor

* add dirs & sign xpu kernel

* add mean cpu&cuda kernel impl

* move sign & mean xpu & npu kernel

* add selected_rows basic impl

* refactor design, BaseTensor to DenseTensor, etc.

* add scale mkldnn kernel

* polish xpu & npu impl details

* fix mkldnn reuse compile failed

* change tensor operation lib name

* rename util filename

* add more comments

* change TensorImplInterface to TensorInterface

* add kernel key and factory

* remove MKLDNNTensorMeta, add MKLDNNDenseTensor

* change XXDeviceContext to XXContext

* add base kernel registrar utils & test on sign

* replace boost::any by paddle::any

* fix several ci failed

* fix npu compile error

* add ordered map util

* fix multiple ordered_map compile errors

* move dev into include dir

* support sign op in static op run

* fix static op run error

* fix new executor compile failed

* add dygraph branch & remove sign_op.h

* fix test_infer_no_need_buffer_slots

* fix rocm compile link error

* fix unitybuild error & clear glog

* fix npu compile failed

* skip quant trans test

* fix part windows compile problem

* fix xpu enforce error

* fix inference test failed

* remove ordered_map to solve quant failed

* fix part of rcom compile faild

* add more register kernels

* revert scale kernel temporarily

* fix code format error

* add new kernel registrar marco

* rename top to tcmpt

* revert xpu, npu, mkldnn impl & remove op def

* add kernel args parse functor to auto parse args

* revert some change & add scale kernels

* add op proto in dygraph kernelcontext building

* polish kernel dispatch logic & nameing rule

* fix scale kernel match error

* fix scale test failed

* add mean API and unittest

* test mean api success

* add branch to solve compiled error

* skip clang format error

* add mean skip rule in op_library

* add dot kernel, api and unittest (#6)

* remove old kernel and add symbol link

* fix dot compiled failed

* add merco for module declare

* fix npu and xpu compile error

* revert sign, mean, scale, dot kernel removing

* add comment for keeping old kernel impl

* fix mutable_data error

* fix bfloat16 conflit

* fix inference undef error

* adapt to msvc compile rules

* polish comment for template inst

* add cmake template instantiation for win

* fix backend to place device id bug

* fix ifdef error

* Op2functor (#7)

* add kernel args maker class

* make args maker non-const

* remove debug log

* modify codes by review options

* split constructPrKernelContext function

* fix output name bug

* fix test_mean_op test_sign_op failed

* fill_any_like kernel refactor (#10)

* fill_any_like kernel refactor

* remove useless code of full_like c++ api

* skip dtype for fill_any_like

* add attrs for kernel key constrcut

* add use_pt_kernel Flags to control whether to use pt kernel (#13)

* add use_pt_kernel Flags to control whether to use pt kernel

* change the default value to true for cheking pt kernels

* fix mutable_data cuda place error

* move high level apis into hapi

* remove selectedrows adapting temporarily

* Support Scalar in Tensor Compute Library (#14)

* fill_any_like kernel refactor

* remove useless code of full_like c++ api

* Support Scalar in Tensor Compute Library

* add scalar in dygraph and static graph mode

* keep the basic type for attr, instead of using scalar for all

* merge the code

* remove mkldnn tensor & polish details

* use flat_hash_map and small_vector in kernel factory

* Refactor flatten kernel (#12)

* refactor flatten kernel

* update infershape function

* fix compile bugs

* fix bugs when merge

* fix compiler bugs

* fix bugs when run test_flatten_api

* fix bugs when run test

* Revert "use flat_hash_map and small_vector in kernel factory"

This reverts commit 23091495cfdd3df8cc1be592d30f09ea66a7c72b.

* Move cpu, cuda and other device code into kernels (#15)

* fill_any_like kernel refactor

* remove useless code of full_like c++ api

* Support Scalar in Tensor Compute Library

* add scalar in dygraph and static graph mode

* keep the basic type for attr, instead of using scalar for all

* merge the code

* start refactor matmul

* move cpu, cuda and other device modules into kernels

* merge code

* polish code in operator.cc

* Perfect unitests (#16)

* perfect unittest

* update license

* replace with flat_hash_map, small_vector (#19)

* fix small_vector build error on windows platform

* replace with flat_hash_map, small_vector

* remove todo

* Perfect unitests (#20)

* perfect unittest

* update license

* fix bug when run tcmpt_utils_test

* refactor execution adapting impl

* fix insert conflit

* Fix CI bug of test_yolov3 (#21)

* fill_any_like kernel refactor

* remove useless code of full_like c++ api

* Support Scalar in Tensor Compute Library

* add scalar in dygraph and static graph mode

* keep the basic type for attr, instead of using scalar for all

* merge the code

* start refactor matmul

* move cpu, cuda and other device modules into kernels

* merge code

* polish code in operator.cc

* Fix CI bug of test_yolov3

* add the tensor base class, test=develop (#17)

* update the tensor base class, test=develop

* remove two funcs, test=develop

* update the error msg, test=develop
Co-authored-by: NChen Weihang <chenweihang@baidu.com>

* [no-verify] commit backend and tensor signature changes

* Rename tcmpt to pten (#23)

* rename tcmpt to pten

* update omitted files for rename to pten

* update omitted file for rename to pten

* remove k of all enum var

* remove kernel_instantiate (#26)

* remove symbols and spatial_tensor

* change common to functions

* readd share tensor impl methods

* add a candidate dense tensor class, test=develop (#28)

* change all Pt to Pten

* resolve conflit with xiaowei

* Op2functor opt1 (#27)

* replace to small vector and change to const &

* add std::move
Co-authored-by: NChen Weihang <chenweihang@baidu.com>

* polish kernel factory and kernel registry

* fix operator test error msg mismatch

* remove tensor signature and backend set member

* move scalar and polish enforce

* revert dtype layout change to fix error

* fix enum operator override error

* add several base unittests

* add pten utils tests

* polish some details

* Dev/op2func refactor 3 (#30)

* add a candidate dense tensor class, test=develop

* remove TensorBase::backend(), test=develop

* remove some ops, test=develop

* cherry-pick the pr of tensor meta, test=develop

* moves the dense tensor and some ops, test=develop

* update the linalg operator, test=develop

* update other operators, test=develop

* fix errors, test=develop

* fix bugs, test=develop

* try to resolve the problem of windows ci, test=develop

* updates codes, test=develop

* fix the tensor_utils.cc, test=develop

* modify the dense tensor, test=develop

* fix the data type, test=develop
Co-authored-by: Nshixiaowei02 <39303645+Shixiaowei02@users.noreply.github.com>

* polish some details

* polish kernel signature details

* fix a bug about offsets of the tensor, test=develop (#31)
Co-authored-by: Nshixiaowei02 <39303645+Shixiaowei02@users.noreply.github.com>

* polish some details
Co-authored-by: Nchentianyu03 <ctychentianyu@gmail.com>
Co-authored-by: Nzyfncg <1370305206@qq.com>
Co-authored-by: NYuanRisheng <yuanrisheng@baidu.com>
Co-authored-by: N石晓伟 <39303645+Shixiaowei02@users.noreply.github.com>

b9fdd3bc

18 9月, 2021 1 次提交

Basic PR on Cost Model (#35774) · 5ba9fe6e

由 Huihuang Zheng 提交于 9月 18, 2021

Add basic Cost Model, it uses executor to run program and profile it to get op time.

This is an early basic version, we will add more functions in the future.

5ba9fe6e

15 9月, 2020 1 次提交
- W
  
  [Pass Compatible] Bind python compatible. (#27262) · f827665a
  由 Wilber 提交于 9月 15, 2020
  
  f827665a
03 6月, 2020 1 次提交

Add crypto python (#24836) · aa47356b

由 Yanghello 提交于 6月 03, 2020

* add crypto helper for paddle, test=develop

* cryptopp.cmake bug fixed, test=develop

* remove debug build type, test=develop

* fixed CMakeLists for new target, test=develop

* fix CI bug, test=develop

* add cmake option flag DWITH_CRYPTO, test=develop

* add crypto api for python, test=develop

* Revert "add crypto api for python, test=develop"

This reverts commit 3a1cfa9d.

* Revert "Add crypto api (#24694)"

This reverts commit 5a7a517c.

* Revert "Revert "Add crypto api (#24694)""

This reverts commit f952b19f.

* fixed cryptopp cmake building error, test=develop

* change WITH_CRYPTO building option to OFF, test=develop

* âfixed cipher test failed, test=develop

* "add crypto api for python, test=develop"

This reverts commit 83fb55c0.

* travis CI bug fixed, test=develop

* fixed test in python3

* test=develop

* fixed unittest, test=develop

aa47356b

21 1月, 2019 1 次提交
- F
  add python inference api (#15248) · d60751fb
  由 flame 提交于 1月 21, 2019
```
add python inference api
```
  d60751fb
10 1月, 2019 1 次提交
- F
  
  Add python ir graph API (#14917) · fb63cd89
  由 flame 提交于 1月 10, 2019
  
  fb63cd89
13 12月, 2018 1 次提交
- S
  fix cmake · deb0d41c
  由 sneaxiy 提交于 12月 12, 2018
```
fix cmake again
test=develop
```
  deb0d41c
10 12月, 2018 1 次提交
- S
  
  featue/py_func · 8760d23c
  由 sneaxiy 提交于 12月 10, 2018
  
  8760d23c
10 9月, 2018 1 次提交
- Y
  
  fea/add color log (#13305) · 2fd1bf2e
  由 Yan Chunwei 提交于 9月 10, 2018
  
  2fd1bf2e
18 6月, 2018 1 次提交
- Y
  
  Feature/pass manager (#11440) · d7345959
  由 Yan Chunwei 提交于 6月 18, 2018
  
  d7345959
24 5月, 2018 1 次提交
- Y
  
  fix inference api (#10867) · b1d44685
  由 Yan Chunwei 提交于 5月 24, 2018
  
  b1d44685
23 5月, 2018 1 次提交
- Y
  Inference analysis/init data flow graph analysis (#10776) · 1153144f
  由 Yan Chunwei 提交于 5月 23, 2018
```
Add the demo of subgraph splitter
```
  1153144f
22 3月, 2018 1 次提交
- Y
  
  Extract SSAGraph · dd73d18b
  由 Yu Yang 提交于 3月 22, 2018
  
  dd73d18b
07 3月, 2018 2 次提交
- Y
  
  Complete RecordIO reader op · 72be7a61
  由 Yu Yang 提交于 3月 07, 2018
  
  72be7a61
- F
  
  fix compile errors · af64f39b
  由 fengjiayi 提交于 3月 07, 2018
  
  af64f39b
06 3月, 2018 2 次提交
- F
  
  init double buffer · 3fcd16ed
  由 fengjiayi 提交于 3月 06, 2018
  
  3fcd16ed
- Y
  
  Extract create_reader_op to three files · 4d8345e3
  由 Yu Yang 提交于 3月 06, 2018
  
  4d8345e3
15 2月, 2018 1 次提交

Update tensor_util.h (#8422) · cfffb1a3

由 Yi Wang 提交于 2月 14, 2018

* Update tensor_util.h

* Update with moved TensorDesc

* Fix tensur_utils.cu

* Update

* Update

* Update

* Update

* Make tensor_util.cu a symbolic link

cfffb1a3

10 2月, 2018 2 次提交
- Y
  
  Correct #include path · fc374821
  由 Yi Wang 提交于 2月 09, 2018
  
  fc374821
- Y
  
  Move file to fluid/; Edit CMakeLists.txt · 90648f33
  由 Yi Wang 提交于 2月 09, 2018
  
  90648f33
07 2月, 2018 1 次提交
- F
  
  fix compile errors · c1349d98
  由 fengjiayi 提交于 2月 07, 2018
  
  c1349d98
06 2月, 2018 2 次提交
- F
  
  refine code and add unit tests · 0bb9c80e
  由 fengjiayi 提交于 2月 06, 2018
  
  0bb9c80e
- F
  
  Add ReadOp · 1010e39b
  由 fengjiayi 提交于 2月 06, 2018
  
  1010e39b
01 2月, 2018 1 次提交
- F
  
  refine inheritance relationship · d8cc21da
  由 fengjiayi 提交于 2月 01, 2018
  
  d8cc21da
31 1月, 2018 1 次提交
- F
  
  draft of Reader classes · f32ca636
  由 fengjiayi 提交于 1月 31, 2018
  
  f32ca636
30 1月, 2018 1 次提交
- F
  
  init reader.h and reader.cc files · 1acad21b
  由 fengjiayi 提交于 1月 30, 2018
  
  1acad21b

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功