Applying LLMs to Active Learning: Towards Cost-Efficient Cross-Task Text Classification without Manually Labeled Data

Zhang, Yejian; Takada, Shingo

Computer Science > Computation and Language

arXiv:2502.16892 (cs)

[Submitted on 24 Feb 2025]

Title:Applying LLMs to Active Learning: Towards Cost-Efficient Cross-Task Text Classification without Manually Labeled Data

Authors:Yejian Zhang, Shingo Takada

View PDF HTML (experimental)

Abstract:Machine learning-based classifiers have been used for text classification, such as sentiment analysis, news classification, and toxic comment classification. However, supervised machine learning models often require large amounts of labeled data for training, and manual annotation is both labor-intensive and requires domain-specific knowledge, leading to relatively high annotation costs. To address this issue, we propose an approach that integrates large language models (LLMs) into an active learning framework. Our approach combines the Robustly Optimized BERT Pretraining Approach (RoBERTa), Generative Pre-trained Transformer (GPT), and active learning, achieving high cross-task text classification performance without the need for any manually labeled data. Furthermore, compared to directly applying GPT for classification tasks, our approach retains over 93% of its classification performance while requiring only approximately 6% of the computational time and monetary cost, effectively balancing performance and resource efficiency. These findings provide new insights into the efficient utilization of LLMs and active learning algorithms in text classification tasks, paving the way for their broader application.

Comments:	Statement in Accordance with IEEE Preprint Policy: This work is intended for submission to the IEEE for possible publication
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2502.16892 [cs.CL]
	(or arXiv:2502.16892v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2502.16892

Submission history

From: Yejian Zhang [view email]
[v1] Mon, 24 Feb 2025 06:43:19 UTC (2,161 KB)

Computer Science > Computation and Language

Title:Applying LLMs to Active Learning: Towards Cost-Efficient Cross-Task Text Classification without Manually Labeled Data

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Applying LLMs to Active Learning: Towards Cost-Efficient Cross-Task Text Classification without Manually Labeled Data

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators