V"Mean"ba: Visual State Space Models only need 1 hidden dimension

Chi, Tien-Yu; Chiang, Hung-Yueh; Chang, Chi-Chih; Huang, Ning-Chi; Wu, Kai-Chiang

Computer Science > Computer Vision and Pattern Recognition

arXiv:2412.16602 (cs)

[Submitted on 21 Dec 2024]

Title:V"Mean"ba: Visual State Space Models only need 1 hidden dimension

Authors:Tien-Yu Chi, Hung-Yueh Chiang, Chi-Chih Chang, Ning-Chi Huang, Kai-Chiang Wu

View PDF HTML (experimental)

Abstract:Vision transformers dominate image processing tasks due to their superior performance. However, the quadratic complexity of self-attention limits the scalability of these systems and their deployment on resource-constrained devices. State Space Models (SSMs) have emerged as a solution by introducing a linear recurrence mechanism, which reduces the complexity of sequence modeling from quadratic to linear. Recently, SSMs have been extended to high-resolution vision tasks. Nonetheless, the linear recurrence mechanism struggles to fully utilize matrix multiplication units on modern hardware, resulting in a computational bottleneck. We address this issue by introducing \textit{VMeanba}, a training-free compression method that eliminates the channel dimension in SSMs using mean operations. Our key observation is that the output activations of SSM blocks exhibit low variances across channels. Our \textit{VMeanba} leverages this property to optimize computation by averaging activation maps across the channel to reduce the computational overhead without compromising accuracy. Evaluations on image classification and semantic segmentation tasks demonstrate that \textit{VMeanba} achieves up to a 1.12x speedup with less than a 3\% accuracy loss. When combined with 40\% unstructured pruning, the accuracy drop remains under 3\%.

Comments:	Accepted by NeurIPS 2024 Machine Learning for Systems workshop
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2412.16602 [cs.CV]
	(or arXiv:2412.16602v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2412.16602

Submission history

From: TienYu Chi [view email]
[v1] Sat, 21 Dec 2024 12:27:07 UTC (3,626 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:V"Mean"ba: Visual State Space Models only need 1 hidden dimension

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:V"Mean"ba: Visual State Space Models only need 1 hidden dimension

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators