FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Zhu, Yuke; Xie, Chi; Liang, Shuang; Zheng, Bo; Guo, Sheng

Computer Science > Computer Vision and Pattern Recognition

arXiv:2411.14228 (cs)

[Submitted on 21 Nov 2024]

Title:FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Authors:Yuke Zhu, Chi Xie, Shuang Liang, Bo Zheng, Sheng Guo

View PDF HTML (experimental)

Abstract:Recent advances on Multi-modal Large Language Models have demonstrated that high-resolution image input is crucial for model capabilities, especially for fine-grained tasks. However, high-resolution images lead to a quadratic increase in the number of visual tokens input into LLMs, resulting in significant computational costs. Current work develop visual token compression methods to achieve efficiency improvements, often at the expense of performance. We argue that removing visual redundancy can simultaneously improve both efficiency and performance. We build a coarse-to-fine visual token compression method, with a vision-guided sampler for compressing redundant regions with low information density, and a text-guided sampler for selecting visual tokens that are strongly correlated with the user this http URL these two modules, the proposed FocusLLaVA achieves improvements in both efficiency and performance. We validate the effectiveness of our approach on a wide range of evaluation datasets.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2411.14228 [cs.CV]
	(or arXiv:2411.14228v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2411.14228

Submission history

From: Yuke Zhu [view email]
[v1] Thu, 21 Nov 2024 15:37:52 UTC (1,252 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators