UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding

Shi, Bowen; Zhao, Peisen; Wang, Zichen; Zhang, Yuhang; Wang, Yaoming; Li, Jin; Dai, Wenrui; Zou, Junni; Xiong, Hongkai; Tian, Qi; Zhang, Xiaopeng

Computer Science > Computer Vision and Pattern Recognition

arXiv:2401.06397v1 (cs)

A newer version of this paper has been withdrawn by Bowen Shi

[Submitted on 12 Jan 2024 (this version), latest version 29 Oct 2024 (v3)]

Title:UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding

Authors:Bowen Shi, Peisen Zhao, Zichen Wang, Yuhang Zhang, Yaoming Wang, Jin Li, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian, Xiaopeng Zhang

View PDF HTML (experimental)

Abstract:Vision-language foundation models, represented by Contrastive language-image pre-training (CLIP), have gained increasing attention for jointly understanding both vision and textual tasks. However, existing approaches primarily focus on training models to match global image representations with textual descriptions, thereby overlooking the critical alignment between local regions and corresponding text tokens. This paper extends CLIP with multi-granularity alignment. Notably, we deliberately construct a new dataset comprising pseudo annotations at various levels of granularities, encompassing image-level, region-level, and pixel-level captions/tags. Accordingly, we develop a unified multi-granularity learning framework, named UMG-CLIP, that simultaneously empowers the model with versatile perception abilities across different levels of detail. Equipped with parameter efficient tuning, UMG-CLIP surpasses current widely used CLIP models and achieves state-of-the-art performance on diverse image understanding benchmarks, including open-world recognition, retrieval, semantic segmentation, and panoptic segmentation tasks. We hope UMG-CLIP can serve as a valuable option for advancing vision-language foundation models.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2401.06397 [cs.CV]
	(or arXiv:2401.06397v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2401.06397

Submission history

From: Bowen Shi [view email]
[v1] Fri, 12 Jan 2024 06:35:09 UTC (4,127 KB)
[v2] Thu, 18 Jan 2024 16:40:22 UTC (1 KB) (withdrawn)
[v3] Tue, 29 Oct 2024 07:05:36 UTC (5,232 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators