TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators

Li, Jianling; Li, Shangzhan; Gao, Zhenye; Shi, Qi; Li, Yuxuan; Wang, Zefan; Huang, Jiacheng; Wang, Haojie; Wang, Jianrong; Han, Xu; Liu, Zhiyuan; Sun, Maosong

Computer Science > Computation and Language

arXiv:2502.14752 (cs)

[Submitted on 20 Feb 2025]

Title:TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators

Authors:Jianling Li, Shangzhan Li, Zhenye Gao, Qi Shi, Yuxuan Li, Zefan Wang, Jiacheng Huang, Haojie Wang, Jianrong Wang, Xu Han, Zhiyuan Liu, Maosong Sun

View PDF HTML (experimental)

Abstract:Triton, a high-level Python-like language designed for building efficient GPU kernels, is widely adopted in deep learning frameworks due to its portability, flexibility, and accessibility. However, programming and parallel optimization still require considerable trial and error from Triton developers. Despite advances in large language models (LLMs) for conventional code generation, these models struggle to generate accurate, performance-optimized Triton code, as they lack awareness of its specifications and the complexities of GPU programming. More critically, there is an urgent need for systematic evaluations tailored to Triton. In this work, we introduce TritonBench, the first comprehensive benchmark for Triton operator generation. TritonBench features two evaluation channels: a curated set of 184 real-world operators from GitHub and a collection of operators aligned with PyTorch interfaces. Unlike conventional code benchmarks prioritizing functional correctness, TritonBench also profiles efficiency performance on widely deployed GPUs aligned with industry applications. Our study reveals that current state-of-the-art code LLMs struggle to generate efficient Triton operators, highlighting a significant gap in high-performance code generation. TritonBench will be available at this https URL.

Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2502.14752 [cs.CL]
	(or arXiv:2502.14752v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2502.14752

Submission history

From: Qi Shi [view email]
[v1] Thu, 20 Feb 2025 17:21:27 UTC (7,285 KB)

Computer Science > Computation and Language

Title:TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators