A New Method to Capturing Compositional Knowledge in Linguistic Space

Wan, Jiahe

Computer Science > Computer Vision and Pattern Recognition

arXiv:2412.15632 (cs)

[Submitted on 20 Dec 2024]

Title:A New Method to Capturing Compositional Knowledge in Linguistic Space

Authors:Jiahe Wan

View PDF HTML (experimental)

Abstract:Compositional understanding allows visual language models to interpret complex relationships between objects, attributes, and relations in images and text. However, most existing methods often rely on hard negative examples and fine-tuning, which can overestimate improvements and are limited by the difficulty of obtaining hard negatives. In this work, we introduce Zero-Shot Compositional Understanding (ZS-CU), a novel task that enhances compositional understanding without requiring hard negative training data. We propose YUKINO (Yielded Compositional Understanding Knowledge via Textual Inversion with NO), which uses textual inversion to map unlabeled images to pseudo-tokens in a pre-trained CLIP model. We propose introducing "no" logical regularization to address the issue of token interaction in inversion. Additionally, we suggest using knowledge distillation to reduce the time complexity of textual inversion. Experimental results show that YUKINO outperforms the existing multi-modal SOTA models by over 8% on the SugarCREPE benchmark, and also achieves significant improvements in image retrieval tasks.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2412.15632 [cs.CV]
	(or arXiv:2412.15632v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2412.15632

Submission history

From: Jiahe Wan [view email]
[v1] Fri, 20 Dec 2024 07:48:09 UTC (2,843 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:A New Method to Capturing Compositional Knowledge in Linguistic Space

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:A New Method to Capturing Compositional Knowledge in Linguistic Space

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators