Egocentric Video-Language Pretraining

Lin, Kevin Qinghong; Wang, Alex Jinpeng; Soldan, Mattia; Wray, Michael; Yan, Rui; Xu, Eric Zhongcong; Gao, Difei; Tu, Rongcheng; Zhao, Wenzhe; Kong, Weijie; Cai, Chengfei; Wang, Hongfa; Damen, Dima; Ghanem, Bernard; Liu, Wei; Shou, Mike Zheng

Computer Science > Computer Vision and Pattern Recognition

arXiv:2206.01670 (cs)

[Submitted on 3 Jun 2022 (v1), last revised 13 Oct 2022 (this version, v2)]

Title:Egocentric Video-Language Pretraining

Authors:Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu, Mike Zheng Shou

View PDF

Abstract:Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person video-text datasets, such as HowTo100M. In this work, we exploit the recently released Ego4D dataset to pioneer Egocentric VLP along three directions. (i) We create EgoClip, a 1st-person video-text pretraining dataset comprising 3.8M clip-text pairs well-chosen from Ego4D, covering a large variety of human daily activities. (ii) We propose a novel pretraining objective, dubbed EgoNCE, which adapts video-text contrastive learning to the egocentric domain by mining egocentric-aware positive and negative samples. (iii) We introduce EgoMCQ, a development benchmark that is close to EgoClip and hence can support effective validation and fast exploration of our design decisions in EgoClip and EgoNCE. Furthermore, we demonstrate strong performance on five egocentric downstream tasks across three datasets: video-text retrieval on EPIC-KITCHENS-100; action recognition on Charades-Ego; natural language query, moment query, and object state change classification on Ego4D challenge benchmarks. The dataset and code are available at this https URL.

Comments:	Accepted by NeurIPS 2022. Double champions at Ego4D and EPIC-Kitchens, CVPR 2022 challenges. 23 pages, 13 figures, 12 tables. Code: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2206.01670 [cs.CV]
	(or arXiv:2206.01670v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2206.01670

Submission history

From: Qinghong Lin [view email]
[v1] Fri, 3 Jun 2022 16:28:58 UTC (8,334 KB)
[v2] Thu, 13 Oct 2022 03:31:05 UTC (8,363 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Egocentric Video-Language Pretraining

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Egocentric Video-Language Pretraining

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators