Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference

Li, Zeping; Yang, Xinlong; Gao, Ziheng; Liu, Ji; Li, Guanchen; Liu, Zhuang; Li, Dong; Peng, Jinzhang; Tian, Lu; Barsoum, Emad

Computer Science > Artificial Intelligence

arXiv:2406.13170 (cs)

[Submitted on 19 Jun 2024 (v1), last revised 18 Oct 2024 (this version, v2)]

Title:Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference

Authors:Zeping Li, Xinlong Yang, Ziheng Gao, Ji Liu, Guanchen Li, Zhuang Liu, Dong Li, Jinzhang Peng, Lu Tian, Emad Barsoum

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) inherently use autoregressive decoding, which lacks parallelism in inference and results in significantly slow inference speed. While methods such as Medusa constructs parallelized heads, they lack adequate information interaction across different prediction positions. To overcome this limitation, we introduce Amphista, an enhanced speculative decoding framework that builds upon Medusa. Specifically, Amphista models an Auto-embedding Block capable of parallel inference, incorporating bi-directional attention to enable interaction between different drafting heads. Additionally, Amphista integrates Staged Adaptation Layers, which ensure a seamless transition of semantic information from the target model's autoregressive inference to the drafting heads' non-autoregressive inference, effectively achieving paradigm shift and feature fusion. Experimental results on Vicuna models using MT-Bench and Spec-Bench demonstrate that Amphista achieves substantial acceleration while maintaining generation quality. On MT-Bench, Amphista delivers up to 2.75$\times$ speedup over vanilla autoregressive decoding and 1.40$\times$ over Medusa on Vicuna 33B in wall-clock time.

Subjects:	Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2406.13170 [cs.AI]
	(or arXiv:2406.13170v2 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2406.13170

Submission history

From: Zeping Li [view email]
[v1] Wed, 19 Jun 2024 02:53:39 UTC (2,173 KB)
[v2] Fri, 18 Oct 2024 04:13:05 UTC (4,686 KB)

Computer Science > Artificial Intelligence

Title:Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators