Computational Bottlenecks of Training Small-scale Large Language Models

Ashkboos, Saleh; Mirzadeh, Iman; Alizadeh, Keivan; Sekhavat, Mohammad Hossein; Nabi, Moin; Farajtabar, Mehrdad; Faghri, Fartash

Computer Science > Machine Learning

arXiv:2410.19456 (cs)

[Submitted on 25 Oct 2024 (v1), last revised 1 Dec 2024 (this version, v2)]

Title:Computational Bottlenecks of Training Small-scale Large Language Models

Authors:Saleh Ashkboos, Iman Mirzadeh, Keivan Alizadeh, Mohammad Hossein Sekhavat, Moin Nabi, Mehrdad Farajtabar, Fartash Faghri

View PDF HTML (experimental)

Abstract:While large language models (LLMs) dominate the AI landscape, Small-scale large Language Models (SLMs) are gaining attention due to cost and efficiency demands from consumers. However, there is limited research on the training behavior and computational requirements of SLMs. In this study, we explore the computational bottlenecks of training SLMs (up to 2B parameters) by examining the effects of various hyperparameters and configurations, including GPU type, batch size, model size, communication protocol, attention type, and the number of GPUs. We assess these factors on popular cloud services using metrics such as loss per dollar and tokens per second. Our findings aim to support the broader adoption and optimization of language model training for low-resource AI research institutes.

Comments:	8 pages, 4 figures
Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2410.19456 [cs.LG]
	(or arXiv:2410.19456v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2410.19456

Submission history

From: Saleh Ashkboos [view email]
[v1] Fri, 25 Oct 2024 10:30:21 UTC (216 KB)
[v2] Sun, 1 Dec 2024 11:27:09 UTC (86 KB)

Computer Science > Machine Learning

Title:Computational Bottlenecks of Training Small-scale Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Computational Bottlenecks of Training Small-scale Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators