Model Parallelism on Distributed Infrastructure: A Literature Review from Theory to LLM Case-Studies

Brakel, Felix; Odyurt, Uraz; Varbanescu, Ana-Lucia

Computer Science > Distributed, Parallel, and Cluster Computing

arXiv:2403.03699 (cs)

[Submitted on 6 Mar 2024]

Title:Model Parallelism on Distributed Infrastructure: A Literature Review from Theory to LLM Case-Studies

Authors:Felix Brakel, Uraz Odyurt, Ana-Lucia Varbanescu

View PDF HTML (experimental)

Abstract:Neural networks have become a cornerstone of machine learning. As the trend for these to get more and more complex continues, so does the underlying hardware and software infrastructure for training and deployment. In this survey we answer three research questions: "What types of model parallelism exist?", "What are the challenges of model parallelism?", and "What is a modern use-case of model parallelism?" We answer the first question by looking at how neural networks can be parallelised and expressing these as operator graphs while exploring the available dimensions. The dimensions along which neural networks can be parallelised are intra-operator and inter-operator. We answer the second question by collecting and listing both implementation challenges for the types of parallelism, as well as the problem of optimally partitioning the operator graph. We answer the last question by collecting and listing how parallelism is applied in modern multi-billion parameter transformer networks, to the extend that this is possible with the limited information shared about these networks.

Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
Cite as:	arXiv:2403.03699 [cs.DC]
	(or arXiv:2403.03699v1 [cs.DC] for this version)
	https://doi.org/10.48550/arXiv.2403.03699

Submission history

From: Uraz Odyurt [view email]
[v1] Wed, 6 Mar 2024 13:29:00 UTC (347 KB)

Computer Science > Distributed, Parallel, and Cluster Computing

Title:Model Parallelism on Distributed Infrastructure: A Literature Review from Theory to LLM Case-Studies

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Distributed, Parallel, and Cluster Computing

Title:Model Parallelism on Distributed Infrastructure: A Literature Review from Theory to LLM Case-Studies

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators