Is it Really Negative? Evaluating Natural Language Video Localization Performance on Multiple Reliable Videos Pool

Yang, Nakyeong; Kim, Minsung; Yoon, Seunghyun; Shin, Joongbo; Jung, Kyomin

Computer Science > Computer Vision and Pattern Recognition

arXiv:2309.16701v2 (cs)

[Submitted on 15 Aug 2023 (v1), revised 18 Mar 2024 (this version, v2), latest version 9 Aug 2024 (v4)]

Title:Is it Really Negative? Evaluating Natural Language Video Localization Performance on Multiple Reliable Videos Pool

Authors:Nakyeong Yang, Minsung Kim, Seunghyun Yoon, Joongbo Shin, Kyomin Jung

View PDF HTML (experimental)

Abstract:With the explosion of multimedia content in recent years, Video Corpus Moment Retrieval (VCMR), which aims to detect a video moment that matches a given natural language query from multiple videos, has become a critical problem. However, existing VCMR studies have a significant limitation since they have regarded all videos not paired with a specific query as negative, neglecting the possibility of including false negatives when constructing the negative video set. In this paper, we propose an MVMR (Massive Videos Moment Retrieval) task that aims to localize video frames within a massive video set, mitigating the possibility of falsely distinguishing positive and negative videos. For this task, we suggest an automatic dataset construction framework by employing textual and visual semantic matching evaluation methods on the existing video moment search datasets and introduce three MVMR datasets. To solve MVMR task, we further propose a strong method, CroCs, which employs cross-directional contrastive learning that selectively identifies the reliable and informative negatives, enhancing the robustness of a model on MVMR task. Experimental results on the introduced datasets reveal that existing video moment search models are easily distracted by negative video frames, whereas our model shows significant performance.

Comments:	15 pages, 10 figures
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2309.16701 [cs.CV]
	(or arXiv:2309.16701v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2309.16701

Submission history

From: Nakyeong Yang [view email]
[v1] Tue, 15 Aug 2023 17:38:55 UTC (2,319 KB)
[v2] Mon, 18 Mar 2024 08:55:36 UTC (11,259 KB)
[v3] Mon, 29 Jul 2024 06:03:24 UTC (1,811 KB)
[v4] Fri, 9 Aug 2024 00:53:10 UTC (1,811 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Is it Really Negative? Evaluating Natural Language Video Localization Performance on Multiple Reliable Videos Pool

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Is it Really Negative? Evaluating Natural Language Video Localization Performance on Multiple Reliable Videos Pool

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators