A Comparative Analysis of Distributed Training Strategies for GPT-2

Patwardhan, Ishan; Gandhi, Shubham; Khare, Om; Joshi, Amit; Sawant, Suraj

Computer Science > Distributed, Parallel, and Cluster Computing

arXiv:2405.15628 (cs)

[Submitted on 24 May 2024]

Title:A Comparative Analysis of Distributed Training Strategies for GPT-2

Authors:Ishan Patwardhan, Shubham Gandhi, Om Khare, Amit Joshi, Suraj Sawant

View PDF HTML (experimental)

Abstract:The rapid advancement in Large Language Models has been met with significant challenges in their training processes, primarily due to their considerable computational and memory demands. This research examines parallelization techniques developed to address these challenges, enabling the efficient and scalable training of Large Language Models. A comprehensive analysis of both data and model parallelism strategies, including Fully Sharded Data Parallelism and Distributed Data-Parallel frameworks, is provided to assess methods that facilitate efficient model training. Furthermore, the architectural complexities and training methodologies of the Generative Pre-Trained Transformer-2 model are explored. The application of these strategies is further investigated, which is crucial in managing the substantial computational and memory demands of training sophisticated models. This analysis not only highlights the effectiveness of these parallel training strategies in enhancing training efficiency but also their role in enabling the scalable training of large language models. Drawing on recent research findings, through a comprehensive literature review, this research underscores the critical role of parallelization techniques in addressing the computational challenges of training state-of-the-art Large Language Models, thereby contributing to the advancement of training more sophisticated and capable artificial intelligence systems.

Comments:	Submitted to the International Journal of Parallel Programming and is currently under review
Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC)
Cite as:	arXiv:2405.15628 [cs.DC]
	(or arXiv:2405.15628v1 [cs.DC] for this version)
	https://doi.org/10.48550/arXiv.2405.15628

Submission history

From: Ishan Patwardhan [view email]
[v1] Fri, 24 May 2024 15:16:50 UTC (4,793 KB)

Computer Science > Distributed, Parallel, and Cluster Computing

Title:A Comparative Analysis of Distributed Training Strategies for GPT-2

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Distributed, Parallel, and Cluster Computing

Title:A Comparative Analysis of Distributed Training Strategies for GPT-2

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators