distributed training
Distributed training is a method of training machine learning models across multiple computational nodes or devices, enabling the handling of larger datasets and more complex models by parallelizing the training process and reducing the time needed for convergence.
- Accelerated Vertical Federated Adversarial Learning through Decoupling Layer-Wise Dependencies
- DUO: No Compromise to Accuracy Degradation
- FlexOLMo: Open Language Models for Flexible Data Use
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
- StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
- StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
- Synergistic Tensor and Pipeline Parallelism
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise