提交 · 04806ffe83529013aa23fd9831b193662e59c7a6 · PaddlePaddle / PaddleDetection

21 1月, 2018 1 次提交
- C
  
  follow comments · 782ddc5f
  由 chengduoZH 提交于 1月 21, 2018
  
  782ddc5f
19 1月, 2018 1 次提交
- C
  
  follow comments · 0468422d
  由 chengduoZH 提交于 1月 19, 2018
  
  0468422d
18 1月, 2018 3 次提交
- C
  
  modify doc · 259858b4
  由 chengduoZH 提交于 1月 18, 2018
  
  259858b4
- C
  
  code refine · 578d60bf
  由 chengduoZH 提交于 1月 18, 2018
  
  578d60bf
- C
  
  add 4-d for matmul_op · 2edc136c
  由 chengduoZH 提交于 1月 18, 2018
  
  2edc136c
26 12月, 2017 1 次提交
- L
  
  unify the indentation of license · 761b3297
  由 Luo Tao 提交于 12月 26, 2017
  
  761b3297
12 12月, 2017 1 次提交

Refine device context (#6433) · 61ec0b95

由 QI JUN 提交于 12月 12, 2017

There are mainly following fixes:

- take `DeviceContext` as the template parameter of math functors and OpKernel instead of `Place`
- remove `eigen_device` interface in base class  `DeviceContext`
- remove `GetEigenDevice` interface in `ExecutionContext` and base class `DeviceContext`
- remove unused `platform::EigenDeviceConverter`
- rename `REGISTER_OP_GPU_KERNEL` to `REGISTER_OP_CUDA_KERNEL`
- rename `USE_GPU_ONLY_OP` to `USE_CUDA_ONLY_OP`

61ec0b95

14 11月, 2017 1 次提交

Fix matmal_op for debug mode · 6a6e4d8d

由 xuwei06 提交于 11月 10, 2017

The dimension is not set correctly and is not being checked in release mode because eigen_assert is not enabled.

6a6e4d8d

11 11月, 2017 1 次提交
- D
  
  Use G++ to compile some cu operators. · f5e36765
  由 dangqingqing 提交于 11月 11, 2017
  
  f5e36765
20 10月, 2017 1 次提交

Remove template parameter for Tensor methods (#4937) · c532b967

由 Yu Yang 提交于 10月 19, 2017

* Remove template parameter for Tensor methods

* Also check the type is correct when data()
* Simplize holder_

* Fix accuracy_op

* Register Code

c532b967

18 10月, 2017 1 次提交

MatMul operator (#4856) · 16489827

由 Markus Kliegl 提交于 10月 17, 2017

* initial matmul operator

Similar to np.matmul, but also has transpose_X and transpose_Y flags,
and only supports tensors from rank 1 to 3 inclusive.

For GPU, uses cublas?gemmStridedBatched. For CPU, uses
cblas_?gemm_batch if available via MKL; otherwise a simple serial
implementation that loops over the batch dimension is employed for now.

16489827

PaddlePaddle / PaddleDetection 大约 1 年 前同步成功

PaddlePaddle / PaddleDetection
大约 1 年前同步成功