提交 · 2a84054372fbc8310cf98b0dedaaad849f78f9c7 · PaddlePaddle / Paddle

19 11月, 2018 1 次提交

Optimize the layer_norm operator with AVX intrinsic function (#14417) · f4c869d8

由 Yihua Xu 提交于 11月 19, 2018

* Optimize layer_norm operator with AVX intrinsic functions

* Revert the wrong modifications

* Implement the jit kernel for layer_norm operator

* Add math headfile to fix the compile issue (test=develop)

* Add math headfile to fix the compile issue (test=develop)

* Fixed the intrinsic headfile issue (test=develop)

* Fix the conflicts (test=develop)

* Revert for CUDA compiler (test=develop)

* Fixed the cuda depency (test=develop)

* Fix the marco issues (test=develop)

f4c869d8

16 11月, 2018 1 次提交

Refine operator cmake (#14413) · a2d9b344

由 Wu Yi 提交于 11月 16, 2018

* wip simplify operator framework

* wip

* wip

* done test=develop

* clean test=develop

* fix test=develop

* fix deps test=develop

* fix cpu build test=develop

* fix tensorrt build test=develop

* fix tests test=develop

* fix test=develop

* fix cpu build test=develop

a2d9b344

04 5月, 2018 1 次提交
- Y
  
  Clean and extract blas · ef6ea790
  由 Yu Yang 提交于 5月 04, 2018
  
  ef6ea790
25 3月, 2018 2 次提交
- X
  
  Pass cpu build · 1a4be55a
  由 Xin Pan 提交于 3月 25, 2018
  
  1a4be55a
- X
  Improve layer_norm speed · 904fa05f
  由 Xin Pan 提交于 3月 25, 2018
```
    transfomer on a single device step time
    reduces from 0.157 to 0.125
```
  904fa05f
15 2月, 2018 1 次提交

Update tensor_util.h (#8422) · cfffb1a3

由 Yi Wang 提交于 2月 14, 2018

* Update tensor_util.h

* Update with moved TensorDesc

* Fix tensur_utils.cu

* Update

* Update

* Update

* Update

* Make tensor_util.cu a symbolic link

cfffb1a3

12 2月, 2018 1 次提交
- Q
  
  Fix the grammar in copyright. (#8403) · 24509f4a
  由 qingqing01 提交于 2月 12, 2018
  
  24509f4a
10 2月, 2018 2 次提交
- Y
  
  Correct #include path · fc374821
  由 Yi Wang 提交于 2月 09, 2018
  
  fc374821
- Y
  
  Move file to fluid/; Edit CMakeLists.txt · 90648f33
  由 Yi Wang 提交于 2月 09, 2018
  
  90648f33
05 2月, 2018 3 次提交
- C
  
  code refine · 67731297
  由 chengduoZH 提交于 2月 05, 2018
  
  67731297
- C
  
  unifid GPU and CPU implementation · df0e74db
  由 chengduoZH 提交于 2月 05, 2018
  
  df0e74db
- C
  
  Separate GPU and CPU implementation · 5092f529
  由 chengduoZH 提交于 2月 03, 2018
  
  5092f529
03 2月, 2018 1 次提交
- C
  
  unifid GPU and CPU implementation · e0333735
  由 chengduoZH 提交于 2月 03, 2018
  
  e0333735
24 1月, 2018 1 次提交
- C
  
  add layer_norm · ca017719
  由 chengduoZH 提交于 1月 22, 2018
  
  ca017719
22 12月, 2017 1 次提交
- Q
  add data layout (#6832) · 6b475981
  由 QI JUN 提交于 12月 22, 2017
```
* add data layout

* fix ci
```
  6b475981
12 12月, 2017 1 次提交

Refine device context (#6433) · 61ec0b95

由 QI JUN 提交于 12月 12, 2017

There are mainly following fixes:

- take `DeviceContext` as the template parameter of math functors and OpKernel instead of `Place`
- remove `eigen_device` interface in base class  `DeviceContext`
- remove `GetEigenDevice` interface in `ExecutionContext` and base class `DeviceContext`
- remove unused `platform::EigenDeviceConverter`
- rename `REGISTER_OP_GPU_KERNEL` to `REGISTER_OP_CUDA_KERNEL`
- rename `USE_GPU_ONLY_OP` to `USE_CUDA_ONLY_OP`

61ec0b95

25 10月, 2017 1 次提交

CPU Batch Norm Op (#4964) · ee998a9c

由 Qiao Longfei 提交于 10月 24, 2017

* init batch norm op

* prepare input output

* compute mean_out var_out save_mean save_var on CPU

* active is test

* use eigen to do computation

* complete batch norm forward

* set default momentum to 0.9

* add batch norm grad op in CPU

* add tensor_format and NHWC support, add python test

* add test training

* add batch norm gradient test

* improve comment, fix foward Python UnitTest

* add gradient test

* fix eigen warning

* follow name style

* fix a bug

* change float to T

* add simple forward test

* test with different place

* add backward test

* refine python test

* remove old python test code

* code clean

* follow code style

* update comment

ee998a9c

10 10月, 2017 1 次提交
- A
  
  Implementing the fill constant op for the executor · 6efacc14
  由 Abhinav Arora 提交于 10月 09, 2017
  
  6efacc14
28 9月, 2017 1 次提交
- Y
  
  Add Skeleton of Double support · 3a5693e0
  由 Yu Yang 提交于 9月 27, 2017
  
  3a5693e0
20 9月, 2017 1 次提交
- D
  
  Share LoD between input and output of each opeators. · b65709e4
  由 dangqingqing 提交于 9月 19, 2017
  
  b65709e4
23 8月, 2017 1 次提交
- D
  
  Remove set functor and add comapre_grad test · f188e22b
  由 dangqingqing 提交于 8月 23, 2017
  
  f188e22b
11 8月, 2017 1 次提交
- Y
  
  Fix python unit tests · c99f84ac
  由 Yu Yang 提交于 8月 11, 2017
  
  c99f84ac
08 8月, 2017 1 次提交
- F
  
  fix bug · 28476676
  由 fengjiayi 提交于 8月 07, 2017
  
  28476676
07 8月, 2017 1 次提交
- D
  
  "remove type alias done." · 72fb86a2
  由 dongzhihong 提交于 8月 07, 2017
  
  72fb86a2
05 8月, 2017 1 次提交
- Y
  
  Reformat paddle/operators/* strictly following Google Style Guide · 9620df44
  由 Yi Wang 提交于 8月 04, 2017
  
  9620df44
02 8月, 2017 1 次提交
- F
  
  Add unittest for `FillZerosLikeOp` · 8bd73159
  由 fengjiayi 提交于 8月 01, 2017
  
  8bd73159
01 8月, 2017 1 次提交
- Y
  
  Follow comments and merge develop · e2fd2bd0
  由 Yu Yang 提交于 8月 01, 2017
  
  e2fd2bd0
26 7月, 2017 1 次提交
- F
  
  Add fill_zeros_like op · a2dc9614
  由 fengjiayi 提交于 7月 26, 2017
  
  a2dc9614
25 7月, 2017 1 次提交
- Y
  Add type_alias to import framework into ops · efc119b4
  由 Yu Yang 提交于 7月 25, 2017
```
Make implement an operator less noisy.
```
  efc119b4
19 7月, 2017 2 次提交
- Q
  
  add Flatten method to EigenVector · d9fa6159
  由 qijun 提交于 7月 19, 2017
  
  d9fa6159
- Y
  
  Update · 00ed5643
  由 Yi Wang 提交于 7月 18, 2017
  
  00ed5643
17 7月, 2017 3 次提交
- Q
  
  set correct place for output tensor · 2a03e380
  由 qijun 提交于 7月 17, 2017
  
  2a03e380
- Y
  Op varient inputs (#2901) · a0caf234
  由 Yan Chunwei 提交于 7月 17, 2017
```
* add inputs

* add ut for multiple inputs

* fix AddToLayer

* op_desc -> op_proto

* CreateArgumentOffsetMap -> CreateInOutOffsetMap

* move CreateInOutOffsetMap from OperatorBase to op registry

* arg_idxs_ -> in_out_idxs_
```
  a0caf234
- Q
  
  implement add_op kernel · d649dbf4
  由 qijun 提交于 7月 17, 2017
  
  d649dbf4
14 7月, 2017 1 次提交
- Q
  
  add_op kernel implementation · bac1426d
  由 qijun 提交于 7月 14, 2017
  
  bac1426d
13 7月, 2017 2 次提交

Follow comments · 79b70c2d

由 Yu Yang 提交于 7月 13, 2017

* Convert `op` --> `operators`
* Remove AddType in OpProtoMaker, because type is part of registry.
* Rename CPU_OR_GPU --> DEVICE_TYPE in registry macro.

79b70c2d

Add a sample op, `add_op` · a0aaafe9

由 Yu Yang 提交于 7月 13, 2017

* Refine register methods, make Op can get rid of whole-archieve
* `USE_OP` before a op is used.
* Add unittest for add_op.

a0aaafe9

PaddlePaddle / Paddle 大约 1 年 前同步成功

PaddlePaddle / Paddle
大约 1 年前同步成功