Towards Understanding Systems Trade-offs in Retrieval-Augmented Generation Model Inference

Shen, Michael; Umar, Muhammad; Maeng, Kiwan; Suh, G. Edward; Gupta, Udit

Computer Science > Hardware Architecture

arXiv:2412.11854 (cs)

[Submitted on 16 Dec 2024]

Title:Towards Understanding Systems Trade-offs in Retrieval-Augmented Generation Model Inference

Authors:Michael Shen, Muhammad Umar, Kiwan Maeng, G. Edward Suh, Udit Gupta

View PDF HTML (experimental)

Abstract:The rapid increase in the number of parameters in large language models (LLMs) has significantly increased the cost involved in fine-tuning and retraining LLMs, a necessity for keeping models up to date and improving accuracy. Retrieval-Augmented Generation (RAG) offers a promising approach to improving the capabilities and accuracy of LLMs without the necessity of retraining. Although RAG eliminates the need for continuous retraining to update model data, it incurs a trade-off in the form of slower model inference times. Resultingly, the use of RAG in enhancing the accuracy and capabilities of LLMs often involves diverse performance implications and trade-offs based on its design. In an effort to begin tackling and mitigating the performance penalties associated with RAG from a systems perspective, this paper introduces a detailed taxonomy and characterization of the different elements within the RAG ecosystem for LLMs that explore trade-offs within latency, throughput, and memory. Our study reveals underlying inefficiencies in RAG for systems deployment, that can result in TTFT latencies that are twice as long and unoptimized datastores that consume terabytes of storage.

Subjects:	Hardware Architecture (cs.AR)
Cite as:	arXiv:2412.11854 [cs.AR]
	(or arXiv:2412.11854v1 [cs.AR] for this version)
	https://doi.org/10.48550/arXiv.2412.11854

Submission history

From: Michael Shen [view email]
[v1] Mon, 16 Dec 2024 15:12:53 UTC (125 KB)

Computer Science > Hardware Architecture

Title:Towards Understanding Systems Trade-offs in Retrieval-Augmented Generation Model Inference

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Hardware Architecture

Title:Towards Understanding Systems Trade-offs in Retrieval-Augmented Generation Model Inference

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators