Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies

Tao, Chaofan; Liu, Qian; Dou, Longxu; Muennighoff, Niklas; Wan, Zhongwei; Luo, Ping; Lin, Min; Wong, Ngai

Computer Science > Computation and Language

arXiv:2407.13623 (cs)

[Submitted on 18 Jul 2024 (v1), last revised 26 Jul 2024 (this version, v2)]

Title:Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies

Authors:Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, Ngai Wong

View PDF HTML (experimental)

Abstract:Research on scaling large language models (LLMs) has primarily focused on model parameters and training data size, overlooking the role of vocabulary size. We investigate how vocabulary size impacts LLM scaling laws by training models ranging from 33M to 3B parameters on up to 500B characters with various vocabulary configurations. We propose three complementary approaches for predicting the compute-optimal vocabulary size: IsoFLOPs analysis, derivative estimation, and parametric fit of the loss function. Our approaches converge on the same result that the optimal vocabulary size depends on the available compute budget and that larger models deserve larger vocabularies. However, most LLMs use too small vocabulary sizes. For example, we predict that the optimal vocabulary size of Llama2-70B should have been at least 216K, 7 times larger than its vocabulary of 32K. We validate our predictions empirically by training models with 3B parameters across different FLOPs budgets. Adopting our predicted optimal vocabulary size consistently improves downstream performance over commonly used vocabulary sizes. By increasing the vocabulary size from the conventional 32K to 43K, we improve performance on ARC-Challenge from 29.1 to 32.0 with the same 2.3e21 FLOPs. Our work emphasizes the necessity of jointly considering model parameters and vocabulary size for efficient scaling.

Comments:	26 pages, 12 figures. Add more related work
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2407.13623 [cs.CL]
	(or arXiv:2407.13623v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2407.13623

Submission history

From: Qian Liu [view email]
[v1] Thu, 18 Jul 2024 15:58:54 UTC (14,847 KB)
[v2] Fri, 26 Jul 2024 12:59:47 UTC (11,063 KB)

Computer Science > Computation and Language

Title:Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators