Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation

Rouzegar, Hamidreza; Makrehchi, Masoud

Computer Science > Computation and Language

arXiv:2406.12114 (cs)

[Submitted on 17 Jun 2024]

Title:Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation

Authors:Hamidreza Rouzegar, Masoud Makrehchi

View PDF HTML (experimental)

Abstract:In the context of text classification, the financial burden of annotation exercises for creating training data is a critical issue. Active learning techniques, particularly those rooted in uncertainty sampling, offer a cost-effective solution by pinpointing the most instructive samples for manual annotation. Similarly, Large Language Models (LLMs) such as GPT-3.5 provide an alternative for automated annotation but come with concerns regarding their reliability. This study introduces a novel methodology that integrates human annotators and LLMs within an Active Learning framework. We conducted evaluations on three public datasets. IMDB for sentiment analysis, a Fake News dataset for authenticity discernment, and a Movie Genres dataset for multi-label classification.The proposed framework integrates human annotation with the output of LLMs, depending on the model uncertainty levels. This strategy achieves an optimal balance between cost efficiency and classification performance. The empirical results show a substantial decrease in the costs associated with data annotation while either maintaining or improving model accuracy.

Comments:	Publisher: Association for Computational Linguistics URL: this https URL
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
MSC classes:	68T50
ACM classes:	I.2.7
Cite as:	arXiv:2406.12114 [cs.CL]
	(or arXiv:2406.12114v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2406.12114
Journal reference:	Proceedings of The 18th Linguistic Annotation Workshop (LAW-XVIII), 2024, pp. 98-111

Submission history

From: Hamidreza Rouzegar [view email]
[v1] Mon, 17 Jun 2024 21:45:48 UTC (8,382 KB)

Computer Science > Computation and Language

Title:Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators