From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs

Ni, Jiliang; Pu, Jiachen; Yang, Zhongyi; Zhou, Kun; Wang, Hui; Xiao, Xiaoliang; Wang, Dakui; Li, Xin; Luo, Jingfeng; Hu, Conggang

Computer Science > Computation and Language

arXiv:2504.13471 (cs)

[Submitted on 18 Apr 2025 (v1), last revised 24 Apr 2025 (this version, v2)]

Title:From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs

Authors:Jiliang Ni, Jiachen Pu, Zhongyi Yang, Kun Zhou, Hui Wang, Xiaoliang Xiao, Dakui Wang, Xin Li, Jingfeng Luo, Conggang Hu

View PDF HTML (experimental)

Abstract:In recent years, Large Language Models (LLMs) have significantly advanced artificial intelligence by optimizing traditional Natural Language Processing (NLP) pipelines, improving performance and generalization. This has spurred their integration into various systems. Many NLP systems, including ours, employ a "one-stage" pipeline directly incorporating LLMs. While effective, this approach incurs substantial costs and latency due to the need for large model parameters to achieve satisfactory outcomes. This paper introduces a three-stage cost-efficient end-to-end LLM deployment pipeline-including prototyping, knowledge transfer, and model compression-to tackle the cost-performance dilemma in LLM-based frameworks. Our approach yields a super tiny model optimized for cost and performance in online systems, simplifying the system architecture. Initially, by transforming complex tasks into a function call-based LLM-driven pipeline, an optimal performance prototype system is constructed to produce high-quality data as a teacher model. The second stage combines techniques like rejection fine-tuning, reinforcement learning, and knowledge distillation to transfer knowledge to a smaller 0.5B student model, delivering effective performance at minimal cost. The final stage applies quantization and pruning to extremely compress models to 0.4B, achieving ultra-low latency and cost. The framework's modular design and cross-domain capabilities suggest potential applicability in other NLP areas.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2504.13471 [cs.CL]
	(or arXiv:2504.13471v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2504.13471

Submission history

From: Jiliang Ni [view email]
[v1] Fri, 18 Apr 2025 05:25:22 UTC (2,188 KB)
[v2] Thu, 24 Apr 2025 07:30:24 UTC (2,914 KB)

Computer Science > Computation and Language

Title:From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators