Automatically Planning Optimal Parallel Strategy for Large Language Models

Li, Zongbiao; Li, Xiezhao; Cui, Yinghao; Chen, Yijun; Gu, Zhixuan; Liu, Yuxuan; Zhu, Wenbo; Jia, Fei; Liu, Ke; Li, Qifeng; Zhan, Junyao; Zhou, Jiangtao; Zhang, Chenxi; Liu, Qike

Computer Science > Artificial Intelligence

arXiv:2501.00254 (cs)

[Submitted on 31 Dec 2024]

Title:Automatically Planning Optimal Parallel Strategy for Large Language Models

Authors:Zongbiao Li (1), Xiezhao Li (1), Yinghao Cui (1), Yijun Chen (1), Zhixuan Gu (1), Yuxuan Liu (1), Wenbo Zhu (1), Fei Jia (1), Ke Liu (1), Qifeng Li (1), Junyao Zhan (1), Jiangtao Zhou (1), Chenxi Zhang (1), Qike Liu (1) ((1) HUAWEI)

View PDF HTML (experimental)

Abstract:The number of parameters in large-scale language models based on transformers is gradually increasing, and the scale of computing clusters is also growing. The technology of quickly mobilizing large amounts of computing resources for parallel computing is becoming increasingly important. In this paper, we propose an automatic parallel algorithm that automatically plans the parallel strategy with maximum throughput based on model and hardware information. By decoupling the training time into computation, communication, and overlap, we established a training duration simulation model. Based on this simulation model, we prune the parallel solution space to shorten the search time required. The multi-node experiment results show that the algorithm can estimate the parallel training duration in real time with an average accuracy of 96%. In our test, the recommendation strategy provided by the algorithm is always globally optimal.

Subjects:	Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2501.00254 [cs.AI]
	(or arXiv:2501.00254v1 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2501.00254

Submission history

From: Xiezhao Li [view email]
[v1] Tue, 31 Dec 2024 03:51:14 UTC (533 KB)

Computer Science > Artificial Intelligence

Title:Automatically Planning Optimal Parallel Strategy for Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:Automatically Planning Optimal Parallel Strategy for Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators