SqueezeSAM: User friendly mobile interactive segmentation

Varadarajan, Balakrishnan; Soran, Bilge; Iandola, Forrest; Xiang, Xiaoyu; Xiong, Yunyang; Wu, Lemeng; Zhu, Chenchen; Krishnamoorthi, Raghuraman; Chandra, Vikas

Computer Science > Computer Vision and Pattern Recognition

arXiv:2312.06736v2 (cs)

[Submitted on 11 Dec 2023 (v1), revised 15 May 2024 (this version, v2), latest version 20 May 2024 (v3)]

Title:SqueezeSAM: User friendly mobile interactive segmentation

Authors:Balakrishnan Varadarajan, Bilge Soran, Forrest Iandola, Xiaoyu Xiang, Yunyang Xiong, Lemeng Wu, Chenchen Zhu, Raghuraman Krishnamoorthi, Vikas Chandra

View PDF HTML (experimental)

Abstract:The Segment Anything Model (SAM) has been a cornerstone in the field of interactive segmentation, propelling significant progress in generative AI, computational photography, and medical imaging. Despite its ability to process arbitrary user input and generate corresponding segmentation masks, SAM's 600 million parameter architecture, based on ViT-H, is not compatible with current mobile hardware due to its high computational demands and large model size. Our research aims to adapt SAM for use in mobile photography applications. To this end, we have developed a fully convolutional SqueezeSAM model architecture, which is 62.5 times faster and 31.6 times smaller than the original SAM, making it a viable solution for mobile applications. Furthermore, our tiny model achieves an mIOU within \emph{1\%} of the original VIT-H architecture.
Automated segmentation holds significant value in the creation flow for photography applications, as evidenced by its adoption by leading industry players like apple and capcut. To facilitate this automation, we employ salient object detection and simulate potential user clicks for foreground object selection, generating an initial segmentation mask that users can subsequently edit interactively. A common user expectation is that a click on a specific part of an object will result in the segmentation of the entire object. For example, a click on a person's t-shirt in a photo should ideally segment the entire person, not just the t-shirt. However, SAM typically only segments the clicked area. We address this limitation through a novel data augmentation scheme. Consequently, if a user clicks on a person holding a basketball, both the person and the basketball are segmented together, aligning with user expectations and enhancing the overall user experience.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2312.06736 [cs.CV]
	(or arXiv:2312.06736v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2312.06736

Submission history

From: Balakrishnan Varadarajan [view email]
[v1] Mon, 11 Dec 2023 16:04:22 UTC (38,254 KB)
[v2] Wed, 15 May 2024 00:40:36 UTC (38,255 KB)
[v3] Mon, 20 May 2024 22:58:52 UTC (38,255 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:SqueezeSAM: User friendly mobile interactive segmentation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:SqueezeSAM: User friendly mobile interactive segmentation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators