1. 01 3月, 2021 1 次提交
    • T
      [Cherry pick] cherry-pick #31102 #30750 #30626 (#31336) · ff4612a3
      Thunderbrook 提交于
      * solve build gpu task core (#30626)
      
      * build gpu task core
      
      * format
      
      * dump to cpu (#30750)
      
      * dump to cpu
      
      * format
      
      * format
      
      * format
      
      * support multi node in heterps (#31102)
      
      * push multi node
      
      * multi node
      
      * MultiThread
      
      * remove log
      
      * solve bug in 30829
      
      * optimizer
      ff4612a3
  2. 29 9月, 2020 1 次提交
  3. 28 9月, 2020 3 次提交
  4. 30 8月, 2020 1 次提交
  5. 06 8月, 2020 1 次提交
    • T
      add heter ps mode (#25682) · 0cb60c70
      Thunderbrook 提交于
      * add heter ps mode
      
      * code style
      test=develop
      
      * add with_pslib
      test=develop
      
      * unitest
      test=develop
      
      * code style
      test=develop
      
      * code style
      test=develop
      
      * code style
      test=develop
      
      * code style
      test=develop
      
      * code style
      test=develop
      
      * code style
      test=develop
      
      * code style
      test=develop
      
      * code style
      test=develop
      
      * test monitor
      test=develop
      
      * prepare trainer
      test=develop
      
      * code style
      test=develop
      0cb60c70
  6. 30 7月, 2020 1 次提交
  7. 04 6月, 2020 1 次提交
  8. 30 4月, 2020 1 次提交
  9. 18 3月, 2020 1 次提交
  10. 17 3月, 2020 1 次提交
  11. 23 2月, 2020 1 次提交
  12. 02 2月, 2020 1 次提交
  13. 20 11月, 2019 1 次提交
  14. 31 10月, 2019 1 次提交
  15. 23 9月, 2019 1 次提交
  16. 06 9月, 2019 1 次提交
  17. 16 8月, 2019 1 次提交
  18. 12 8月, 2019 1 次提交
  19. 25 7月, 2019 1 次提交
  20. 22 7月, 2019 1 次提交
  21. 10 7月, 2019 1 次提交
  22. 08 7月, 2019 1 次提交
  23. 02 7月, 2019 1 次提交
  24. 27 6月, 2019 1 次提交
    • H
      supports collective communicated training (#18175) · b7128bac
      HaoRen 提交于
      * fix prepare context redundant code problem, optimize executor by caching create_varaiables
      test=develop
      
      * supports collective training in executor
      
      * make fetch_list runable with variables, add more unittest for use_program_cache
      test=develop
      
      * fix comment
      test=develop
      
      * use unique name for nccl_id
      
      * supports output to stream in program_to_code
      
      * insert sync_comm_stream before regularization; add skip_op_callstack capability in program_to_code
      
      * set op role in collective training
      
      * add collective op role
      
      * remove orig file
      
      * add build optimizer by strategy
      
      * add collective strategy
      
      * refine collective strategy
      
      * add multi-process role maker
      
      * refine strategy building factory so that we can easily plugin more strategy
      
      * scale loss grad in collective sgd transpiler
      
      * add support for distributed fc
      
      * code format
      
      * revert some features for dist fc
      
      * add support for distributed fc training
      
      * fix prepare context redundant code problem, optimize executor by caching create_varaiables
      test=develop
      
      * supports collective training in executor
      
      * make fetch_list runable with variables, add more unittest for use_program_cache
      test=develop
      
      * use unique name for nccl_id
      
      * supports output to stream in program_to_code
      
      * insert sync_comm_stream before regularization; add skip_op_callstack capability in program_to_code
      
      * set op role in collective training
      
      * add collective op role
      
      * fix comment
      test=develop
      
      * remove orig file
      
      * add build optimizer by strategy
      
      * add collective strategy
      
      * refine collective strategy
      
      * add multi-process role maker
      
      * refine strategy building factory so that we can easily plugin more strategy
      
      * scale loss grad in collective sgd transpiler
      
      * add support for distributed fc
      
      * code format
      
      * revert some features for dist fc
      
      * add support for distributed fc training
      
      * test=develop
      add collective op unittest standard
      
      * test=develop
      remove the test_collective directory
      
      * test=develop
      remove the test_collective directory
      
      * remove slicegather test
      
      * code format for reducescatter
      
      * update attr of shard_index_op
      
      * Modify macro nccl_helper
      
      * remove test without distribute
      
      * macro collective_helper
      
      * marcro update
      
      * test=develop
      update support python3.5
      
      * test=develop change gpu memory use to 0.1 when test
      
      * test=develop
      update ut equal func
      
      * test=develop
      set flags to 1.5
      
      * test=develop fix pickle dumple  py35
      
      * test=develop
      fix divide in slice and add sync_comm_stream
      update atol and rtol to 1e-05
      rm shard_index op and test
      modify read input from file to read from memory
      remove origin_program in framework and add i/o in c_sync_calc_stream
      
      * test=develop update unittest sync operator I/O
      b7128bac
  25. 23 6月, 2019 1 次提交
  26. 17 6月, 2019 1 次提交
  27. 12 6月, 2019 1 次提交
  28. 11 6月, 2019 1 次提交
  29. 23 5月, 2019 1 次提交
  30. 15 5月, 2019 1 次提交
    • J
      support config file, cvm, load, save, shrink (#17319) · 34369944
      jiaqi 提交于
      * support config file, cvm, load, save, shrink
      test=develop
      
      * fix error of worker_num & add table.compress_in_save
      test=develop
      
      * fix code style
      test=develop
      
      * fix save model bug
      test=develop
      34369944
  31. 09 5月, 2019 1 次提交
  32. 25 4月, 2019 1 次提交
  33. 11 4月, 2019 1 次提交
  34. 09 4月, 2019 1 次提交
  35. 30 3月, 2019 1 次提交
  36. 29 3月, 2019 3 次提交