What's in a Name? Beyond Class Indices for Image Recognition

Han, Kai; Li, Yandong; Vaze, Sagar; Li, Jie; Jia, Xuhui

Computer Science > Computer Vision and Pattern Recognition

arXiv:2304.02364v1 (cs)

[Submitted on 5 Apr 2023 (this version), latest version 27 Jul 2024 (v2)]

Title:What's in a Name? Beyond Class Indices for Image Recognition

Authors:Kai Han, Yandong Li, Sagar Vaze, Jie Li, Xuhui Jia

View PDF

Abstract:Existing machine learning models demonstrate excellent performance in image object recognition after training on a large-scale dataset under full supervision. However, these models only learn to map an image to a predefined class index, without revealing the actual semantic meaning of the object in the image. In contrast, vision-language models like CLIP are able to assign semantic class names to unseen objects in a `zero-shot' manner, although they still rely on a predefined set of candidate names at test time. In this paper, we reconsider the recognition problem and task a vision-language model to assign class names to images given only a large and essentially unconstrained vocabulary of categories as prior information. We use non-parametric methods to establish relationships between images which allow the model to automatically narrow down the set of possible candidate names. Specifically, we propose iteratively clustering the data and voting on class names within them, showing that this enables a roughly 50\% improvement over the baseline on ImageNet. Furthermore, we tackle this problem both in unsupervised and partially supervised settings, as well as with a coarse-grained and fine-grained search space as the unconstrained dictionary.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2304.02364 [cs.CV]
	(or arXiv:2304.02364v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2304.02364

Submission history

From: Kai Han [view email]
[v1] Wed, 5 Apr 2023 11:01:23 UTC (6,800 KB)
[v2] Sat, 27 Jul 2024 15:07:38 UTC (5,013 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:What's in a Name? Beyond Class Indices for Image Recognition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:What's in a Name? Beyond Class Indices for Image Recognition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators