Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation

Ahn, Jinwoo; Kwon, Hyeokjoon; Yoo, Hwiyeon

Computer Science > Computer Vision and Pattern Recognition

arXiv:2411.15620 (cs)

[Submitted on 23 Nov 2024]

Title:Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation

Authors:Jinwoo Ahn, Hyeokjoon Kwon, Hwiyeon Yoo

View PDF HTML (experimental)

Abstract:Recent advent of vision-based foundation models has enabled efficient and high-quality object detection at ease. Despite the success of previous studies, object detection models face limitations on capturing small components from holistic objects and taking user intention into account. To address these challenges, we propose a novel foundation model-based detection method called FOCUS: Fine-grained Open-Vocabulary Object ReCognition via User-Guided Segmentation. FOCUS merges the capabilities of vision foundation models to automate open-vocabulary object detection at flexible granularity and allow users to directly guide the detection process via natural language. It not only excels at identifying and locating granular constituent elements but also minimizes unnecessary user intervention yet grants them significant control. With FOCUS, users can make explainable requests to actively guide the detection process in the intended direction. Our results show that FOCUS effectively enhances the detection capabilities of baseline models and shows consistent performance across varying object types.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2411.15620 [cs.CV]
	(or arXiv:2411.15620v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2411.15620

Submission history

From: Jinwoo Ahn [view email]
[v1] Sat, 23 Nov 2024 18:13:27 UTC (3,471 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators