STF: Sentence Transformer Fine-Tuning For Topic Categorization With Limited Data

Daouadi, Kheir Eddine; Boualleg, Yaakoub; Guehairia, Oussama

Computer Science > Computation and Language

arXiv:2407.03253 (cs)

[Submitted on 3 Jul 2024]

Title:STF: Sentence Transformer Fine-Tuning For Topic Categorization With Limited Data

Authors:Kheir Eddine Daouadi, Yaakoub Boualleg, Oussama Guehairia

View PDF HTML (experimental)

Abstract:Nowadays, topic classification from tweets attracts considerable research attention. Different classification systems have been suggested thanks to these research efforts. Nevertheless, they face major challenges owing to low performance metrics due to the limited amount of labeled data. We propose Sentence Transformers Fine-tuning (STF), a topic detection system that leverages pretrained Sentence Transformers models and fine-tuning to classify topics from tweets accurately. Moreover, extensive parameter sensitivity analyses were conducted to finetune STF parameters for our topic classification task to achieve the best performance results. Experiments on two benchmark datasets demonstrated that (1) the proposed STF can be effectively used for classifying tweet topics and outperforms the latest state-of-the-art approaches, and (2) the proposed STF does not require a huge amount of labeled tweets to achieve good accuracy, which is a limitation of many state-of-the-art approaches. Our main contribution is the achievement of promising results in tweet topic classification by applying pretrained sentence transformers language models.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2407.03253 [cs.CL]
	(or arXiv:2407.03253v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2407.03253

Submission history

From: Kheir Eddine Daouadi [view email]
[v1] Wed, 3 Jul 2024 16:34:56 UTC (20 KB)

Computer Science > Computation and Language

Title:STF: Sentence Transformer Fine-Tuning For Topic Categorization With Limited Data

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:STF: Sentence Transformer Fine-Tuning For Topic Categorization With Limited Data

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators