Learning Representations from Audio-Visual Spatial Alignment

Morgado, Pedro; Li, Yi; Vasconcelos, Nuno

Computer Science > Computer Vision and Pattern Recognition

arXiv:2011.01819 (cs)

[Submitted on 3 Nov 2020]

Title:Learning Representations from Audio-Visual Spatial Alignment

Authors:Pedro Morgado, Yi Li, Nuno Vasconcelos

View PDF

Abstract:We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based on audio-visual correspondence (AVC) predict whether audio and video clips originate from the same or different video instances. Audio-visual temporal synchronization (AVTS) further discriminates negative pairs originated from the same video instance but at different moments in time. While these approaches learn high-quality representations for downstream tasks such as action recognition, their training objectives disregard spatial cues naturally occurring in audio and visual signals. To learn from these spatial cues, we tasked a network to perform contrastive audio-visual spatial alignment of 360° video and spatial audio. The ability to perform spatial alignment is enhanced by reasoning over the full spatial content of the 360° video using a transformer architecture to combine representations from multiple viewpoints. The advantages of the proposed pretext task are demonstrated on a variety of audio and visual downstream tasks, including audio-visual correspondence, spatial alignment, action recognition, and video semantic segmentation.

Comments:	To appear at Advances in Neural Information Processing Systems (NeurIPS), 2020
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2011.01819 [cs.CV]
	(or arXiv:2011.01819v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2011.01819

Submission history

From: Pedro Morgado [view email]
[v1] Tue, 3 Nov 2020 16:20:04 UTC (1,378 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Learning Representations from Audio-Visual Spatial Alignment

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Learning Representations from Audio-Visual Spatial Alignment

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators