Localizing Events in Videos with Multimodal Queries

Zhang, Gengyuan; Fok, Mang Ling Ada; Xia, Yan; Tang, Yansong; Cremers, Daniel; Torr, Philip; Tresp, Volker; Gu, Jindong

Computer Science > Computer Vision and Pattern Recognition

arXiv:2406.10079 (cs)

[Submitted on 14 Jun 2024 (v1), last revised 22 Jun 2024 (this version, v2)]

Title:Localizing Events in Videos with Multimodal Queries

Authors:Gengyuan Zhang, Mang Ling Ada Fok, Yan Xia, Yansong Tang, Daniel Cremers, Philip Torr, Volker Tresp, Jindong Gu

View PDF HTML (experimental)

Abstract:Video understanding is a pivotal task in the digital era, yet the dynamic and multievent nature of videos makes them labor-intensive and computationally demanding to process. Thus, localizing a specific event given a semantic query has gained importance in both user-oriented applications like video search and academic research into video foundation models. A significant limitation in current research is that semantic queries are typically in natural language that depicts the semantics of the target event. This setting overlooks the potential for multimodal semantic queries composed of images and texts. To address this gap, we introduce a new benchmark, ICQ, for localizing events in videos with multimodal queries, along with a new evaluation dataset ICQ-Highlight. Our new benchmark aims to evaluate how well models can localize an event given a multimodal semantic query that consists of a reference image, which depicts the event, and a refinement text to adjust the images' semantics. To systematically benchmark model performance, we include 4 styles of reference images and 5 types of refinement texts, allowing us to explore model performance across different domains. We propose 3 adaptation methods that tailor existing models to our new setting and evaluate 10 SOTA models, ranging from specialized to large-scale foundation models. We believe this benchmark is an initial step toward investigating multimodal queries in video event localization.

Comments:	9 pages; fix some typos
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2406.10079 [cs.CV]
	(or arXiv:2406.10079v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2406.10079

Submission history

From: Gengyuan Zhang [view email]
[v1] Fri, 14 Jun 2024 14:35:58 UTC (8,389 KB)
[v2] Sat, 22 Jun 2024 06:53:40 UTC (9,049 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Localizing Events in Videos with Multimodal Queries

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Localizing Events in Videos with Multimodal Queries

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators