Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling

September 05, 2024 · Declared Dead · 🏛 International Conference on Architectural Support for Programming Languages and Operating Systems

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Yujie Wang, Shenhan Zhu, Fangcheng Fu, Xupeng Miao, Jie Zhang, Juan Zhu, Fan Hong, Yong Li, Bin Cui arXiv ID 2409.03365 Category cs.DC: Distributed Computing Cross-listed cs.LG Citations 6 Venue International Conference on Architectural Support for Programming Languages and Operating Systems Last Checked 3 months ago

Abstract

Recent foundation models are capable of handling multiple tasks and multiple data modalities with the unified base model structure and several specialized model components. However, efficient training of such multi-task (MT) multi-modal (MM) models poses significant system challenges due to the sophisticated model architecture and the heterogeneous workloads of different tasks and modalities. In this paper, we propose Spindle, a brand new training system tailored for resource-efficient and high-performance training of MT MM models via wavefront scheduling. The key idea of Spindle is to decompose the model execution into waves and address the joint optimization problem sequentially, including both heterogeneity-aware workload parallelization and dependency-driven execution scheduling. We build our system and evaluate it on various MT MM models. Experiments demonstrate the superior performance and efficiency of Spindle, with speedup ratio up to 71% compared to state-of-the-art training systems.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Distributed Computing

R.I.P. 👻 Ghosted

TensorFlow: A system for large-scale machine learning

Martín Abadi, Paul Barham, ... (+20 more)

cs.DC 🏛 OSDI 📚 19.3K cites 10 years ago

R.I.P. 👻 Ghosted

TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems

Martín Abadi, Ashish Agarwal, ... (+38 more)

cs.DC 🏛 arXiv 📚 11.6K cites 10 years ago

R.I.P. 👻 Ghosted

Hyperledger Fabric: A Distributed Operating System for Permissioned Blockchains

Elli Androulaki, Artem Barger, ... (+19 more)

cs.DC 🏛 European Conference on Computer Systems 📚 4.0K cites 8 years ago

R.I.P. 👻 Ghosted

Reproducing GW150914: the first observation of gravitational waves from a binary black hole merger

Duncan A. Brown, Karan Vahi, ... (+3 more)

cs.DC 🏛 Computing in science & engineering (Print) 📚 2.3K cites 5 years ago

R.I.P. 👻 Ghosted

MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems

Tianqi Chen, Mu Li, ... (+8 more)

cs.DC 🏛 arXiv 📚 2.3K cites 10 years ago

R.I.P. 👻 Ghosted

Efficient Architecture-Aware Acceleration of BWA-MEM for Multicore Systems

Vasimuddin Md, Sanchit Misra, ... (+2 more)

cs.DC 🏛 IPDPS 📚 2.3K cites 6 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Language Models are Few-Shot Learners

Tom B. Brown, Benjamin Mann, ... (+29 more)

cs.CL 🏛 NeurIPS 📚 54.2K cites 6 years ago

R.I.P. 👻 Ghosted

PyTorch: An Imperative Style, High-Performance Deep Learning Library

Adam Paszke, Sam Gross, ... (+19 more)

cs.LG 🏛 NeurIPS 📚 49.7K cites 6 years ago

R.I.P. 👻 Ghosted

XGBoost: A Scalable Tree Boosting System

Tianqi Chen, Carlos Guestrin

cs.LG 🏛 KDD 📚 49.2K cites 10 years ago

R.I.P. 👻 Ghosted

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Sergey Ioffe, Christian Szegedy

cs.LG 🏛 ICML 📚 46.0K cites 11 years ago