HindiLLM: Large Language Model for Hindi

Chouhan, Sanjay; Nath, Shubha Brata; Dutta, Aparajita

doi:10.1007/978-3-031-78172-8_17

Computer Science > Computation and Language

arXiv:2412.20357 (cs)

[Submitted on 29 Dec 2024]

Title:HindiLLM: Large Language Model for Hindi

Authors:Sanjay Chouhan, Shubha Brata Nath, Aparajita Dutta

View PDF HTML (experimental)

Abstract:The advancements in the Large Language Model (LLM) have helped in solving several problems related to language processing. Most of the researches have focused on the English language only, because of its popularity and abundance on the internet. However, a high-performance language model for Hindi and other Indic languages is lacking in the literature. In this work, we have pre-trained two autoregressive LLM models for the Hindi language, namely HindiLLM-Small and HindiLLM-Medium. We use a two-step process comprising unsupervised pre-training and supervised fine-tuning. First, we create a large and high-quality text corpus for unsupervised pre-training. Next, we train a Byte-Pair Encoding, named HindiLLM tokenizer, using the pre-training text data. We then perform training on the unlabeled data, known as the pre-training step, to get the HindiLLM base models. Furthermore, we perform fine-tuning of the HindiLLM base models for different tasks like sentiment analysis, text classification, natural language inference, and multiple choice question-answer on popular labeled datasets to measure the real-world performance. The evaluation shows that the HindiLLM-based fine-tuned models outperform several models in most of the language related tasks.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2412.20357 [cs.CL]
	(or arXiv:2412.20357v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2412.20357
Journal reference:	Pattern Recognition. ICPR 2024. Lecture Notes in Computer Science, vol 15306, pp. 255--270, Springer, Cham, 2025
Related DOI:	https://doi.org/10.1007/978-3-031-78172-8_17

Submission history

From: Sanjay Chouhan [view email]
[v1] Sun, 29 Dec 2024 05:28:15 UTC (76 KB)

Computer Science > Computation and Language

Title:HindiLLM: Large Language Model for Hindi

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:HindiLLM: Large Language Model for Hindi

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators