- 02 9月, 2020 1 次提交
-
-
由 Jeff Rasley 提交于
* Sparse attn + ops/runtime refactor + v0.3.0 Co-authored-by: NArash Ashari <arashari@microsoft.com> Co-authored-by: NArash Ashari <arashari@microsoft.com>
-
- 01 9月, 2020 1 次提交
-
-
由 Samyam Rajbhandari 提交于
* Adding gradient accumulation support for ZeRO Stage 2. Changing all Megatron-LM tests to also test gradient accumulation * Gradient Accumulation support for Stage 2. Model tests added to test the feature * formatting * Update deepspeed_light.py removing comment * Update ds_config_func_bs8_zero1.json reverting this file back. Its not needed for this PR * defining baseline prefix Co-authored-by: NJeff Rasley <jerasley@microsoft.com>
-
- 14 7月, 2020 1 次提交
-
-
由 Olatunji Ruwase 提交于
* Support saving and loading ZeRO checkpoints on different data parallelism degree. * Fix formatting * Support checkpoint with varying GPU count in ZeRO stage 1 * Fix formatting * Formatting fixes * Update model tests * Remove pprint * Minor fix * Fix formatting * Update model tests Co-authored-by: NJeff Rasley <jerasley@microsoft.com>
-
- 07 7月, 2020 1 次提交
-
-
由 Olatunji Ruwase 提交于
* Load non-DeepSpeed checkpoints into ZeRO optimizer * Handle parameters smaller than DP * Formatting fixes * Handle empty partitions * Fix perf bug Co-authored-by: NJeff Rasley <jerasley@microsoft.com>
-
- 24 6月, 2020 1 次提交
-
-
由 Olatunji Ruwase 提交于
* Load non-DeepSpeed checkpoints into ZeRO optimizer * Handle parameters smaller than DP * Formatting fixes
-
- 20 6月, 2020 1 次提交
-
-
由 Samyam Rajbhandari 提交于
* Removing handle_overflow debugging code in deepspeed_utils.py * Removing handle_overflow debugging code in deepspeed_zero_optimizer.py Removing unnecessary overflow handle code. Not sure why it was there in the first place.
-
- 05 6月, 2020 1 次提交
-
-
由 Chunyang Wen 提交于
* Add log util * replace all occurrences of print and logging * address format * disable propagate to avoid duplicate log
-
- 04 6月, 2020 1 次提交
-
-
由 eltonzheng 提交于
-
- 28 5月, 2020 2 次提交
-
-
由 Jeff Rasley 提交于
* add support for predivide as a flag * add predivide json config, remove allgather_disable (as it's not currently used anymore)
-
由 Samyam Rajbhandari 提交于
* Fix for CPU memory Bloating Issue caused by pyorch backward graph creation in allgather. Fixed by calling detach on tensors before calling all_gather * Fix for CPU memory Bloating Issue caused by pyorch backward graph creation in allgather. Fixed by calling detach on tensors before calling all_gather * Fix for CPU memory Bloating Issue caused by pyorch backward graph creation in allgather. Fixed by calling detach on tensors before calling all_gather
-
- 19 5月, 2020 1 次提交
-
-
由 Jeff Rasley 提交于
Updates for ZeRO stage 2 + ZeRO stage 1 w. RS Co-authored-by: NTunji Ruwase <olruwase@microsoft.com> Co-authored-by: NSamyam Rajbhandari <samyamr@microsoft.com> Co-authored-by: NShaden Smith <ShadenTSmith@gmail.com> Co-authored-by: NElton Zheng <eltonz@microsoft.com> Co-authored-by: NShaden Smith <Shaden.Smith@microsoft.com> Co-authored-by: Nyuxionghe <yuxhe@microsoft.com> Co-authored-by: NArash Ashari <arashari@microsoft.com>
-
- 25 4月, 2020 1 次提交
-
-
由 Olatunji Ruwase 提交于
-
- 21 4月, 2020 1 次提交
-
-
由 Olatunji Ruwase 提交于
Co-authored-by: NShaden Smith <Shaden.Smith@microsoft.com>
-
- 03 4月, 2020 1 次提交
-
-
由 kouml 提交于
-
- 26 3月, 2020 1 次提交
-
-
由 Shaden Smith 提交于
-
- 11 3月, 2020 1 次提交
-
-
由 Samyam Rajbhandari 提交于
* Enhancement: Ability to load checkpoint without loading the optimizer states. Unittest testing saving and loading checkpoint with fused, unfused and zero optimizer. The unitest takes about 165s
-
- 04 2月, 2020 1 次提交
-
-
由 Samyam Rajbhandari 提交于
Different Optimizers in DeepSpeed.
-