Retrieval-Enhanced Contrastive Vision-Text Models

Iscen, Ahmet; Caron, Mathilde; Fathi, Alireza; Schmid, Cordelia

Computer Science > Computer Vision and Pattern Recognition

arXiv:2306.07196 (cs)

[Submitted on 12 Jun 2023 (v1), last revised 21 Feb 2024 (this version, v2)]

Title:Retrieval-Enhanced Contrastive Vision-Text Models

Authors:Ahmet Iscen, Mathilde Caron, Alireza Fathi, Cordelia Schmid

View PDF HTML (experimental)

Abstract:Contrastive image-text models such as CLIP form the building blocks of many state-of-the-art systems. While they excel at recognizing common generic concepts, they still struggle on fine-grained entities which are rare, or even absent from the pre-training dataset. Hence, a key ingredient to their success has been the use of large-scale curated pre-training data aiming at expanding the set of concepts that they can memorize during the pre-training stage. In this work, we explore an alternative to encoding fine-grained knowledge directly into the model's parameters: we instead train the model to retrieve this knowledge from an external memory. Specifically, we propose to equip existing vision-text models with the ability to refine their embedding with cross-modal retrieved information from a memory at inference time, which greatly improves their zero-shot predictions. Remarkably, we show that this can be done with a light-weight, single-layer, fusion transformer on top of a frozen CLIP. Our experiments validate that our retrieval-enhanced contrastive (RECO) training improves CLIP performance substantially on several challenging fine-grained tasks: for example +10.9 on Stanford Cars, +10.2 on CUB-2011 and +7.3 on the recent OVEN benchmark, where we even outperform the fine-tuned models on unseen classes.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2306.07196 [cs.CV]
	(or arXiv:2306.07196v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2306.07196

Submission history

From: Ahmet Iscen [view email]
[v1] Mon, 12 Jun 2023 15:52:02 UTC (9,535 KB)
[v2] Wed, 21 Feb 2024 16:55:00 UTC (9,755 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Retrieval-Enhanced Contrastive Vision-Text Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Retrieval-Enhanced Contrastive Vision-Text Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators