Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Dong, Juechu; Feng, Boyuan; Guessous, Driss; Liang, Yanbo; He, Horace

Computer Science > Machine Learning

arXiv:2412.05496 (cs)

[Submitted on 7 Dec 2024]

Title:Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Authors:Juechu Dong, Boyuan Feng, Driss Guessous, Yanbo Liang, Horace He

View PDF HTML (experimental)

Abstract:Over the past 7 years, attention has become one of the most important primitives in deep learning. The primary approach to optimize attention is FlashAttention, which fuses the operation together, drastically improving both the runtime and the memory consumption. However, the importance of FlashAttention combined with its monolithic nature poses a problem for researchers aiming to try new attention variants -- a "software lottery". This problem is exacerbated by the difficulty of writing efficient fused attention kernels, resisting traditional compiler-based approaches. We introduce FlexAttention, a novel compiler-driven programming model that allows implementing the majority of attention variants in a few lines of idiomatic PyTorch code. We demonstrate that many existing attention variants (e.g. Alibi, Document Masking, PagedAttention, etc.) can be implemented via FlexAttention, and that we achieve competitive performance compared to these handwritten kernels. Finally, we demonstrate how FlexAttention allows for easy composition of attention variants, solving the combinatorial explosion of attention variants.

Subjects:	Machine Learning (cs.LG); Performance (cs.PF); Programming Languages (cs.PL)
Cite as:	arXiv:2412.05496 [cs.LG]
	(or arXiv:2412.05496v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2412.05496

Submission history

From: Boyuan Feng [view email]
[v1] Sat, 7 Dec 2024 01:46:38 UTC (2,891 KB)

Computer Science > Machine Learning

Title:Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators