Online Scheduling for LLM Inference with KV Cache Constraints

Jaillet, Patrick; Jiang, Jiashuo; Podimata, Chara; Zhou, Zijie

Computer Science > Machine Learning

arXiv:2502.07115 (cs)

[Submitted on 10 Feb 2025 (v1), last revised 13 Feb 2025 (this version, v2)]

Title:Online Scheduling for LLM Inference with KV Cache Constraints

Authors:Patrick Jaillet, Jiashuo Jiang, Chara Podimata, Zijie Zhou

View PDF HTML (experimental)

Abstract:Large Language Model (LLM) inference, where a trained model generates text one word at a time in response to user prompts, is a computationally intensive process requiring efficient scheduling to optimize latency and resource utilization. A key challenge in LLM inference is the management of the Key-Value (KV) cache, which reduces redundant computations but introduces memory constraints. In this work, we model LLM inference with KV cache constraints theoretically and propose novel batching and scheduling algorithms that minimize inference latency while effectively managing the KV cache's memory.
We analyze both semi-online and fully online scheduling models, and our results are threefold. First, we provide a polynomial-time algorithm that achieves exact optimality in terms of average latency in the semi-online prompt arrival model. Second, in the fully online case with a stochastic prompt arrival, we introduce an efficient online scheduling algorithm with constant regret. Third, we prove that no algorithm (deterministic or randomized) can achieve a constant competitive ratio in fully online adversarial settings. Our empirical evaluations on a public LLM inference dataset, using the Llama-70B model on A100 GPUs, show that our approach significantly outperforms benchmark algorithms used currently in practice, achieving lower latency while reducing energy consumption. Overall, our results offer a path toward more sustainable and cost-effective LLM deployment.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
Cite as:	arXiv:2502.07115 [cs.LG]
	(or arXiv:2502.07115v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2502.07115

Submission history

From: Zijie Zhou [view email]
[v1] Mon, 10 Feb 2025 23:11:44 UTC (5,578 KB)
[v2] Thu, 13 Feb 2025 12:54:36 UTC (5,578 KB)

Computer Science > Machine Learning

Title:Online Scheduling for LLM Inference with KV Cache Constraints

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Online Scheduling for LLM Inference with KV Cache Constraints

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators