Visual Representation Learning with Stochastic Frame Prediction

Jang, Huiwon; Kim, Dongyoung; Kim, Junsu; Shin, Jinwoo; Abbeel, Pieter; Seo, Younggyo

Computer Science > Computer Vision and Pattern Recognition

arXiv:2406.07398 (cs)

[Submitted on 11 Jun 2024 (v1), last revised 8 Aug 2024 (this version, v2)]

Title:Visual Representation Learning with Stochastic Frame Prediction

Authors:Huiwon Jang, Dongyoung Kim, Junsu Kim, Jinwoo Shin, Pieter Abbeel, Younggyo Seo

View PDF HTML (experimental)

Abstract:Self-supervised learning of image representations by predicting future frames is a promising direction but still remains a challenge. This is because of the under-determined nature of frame prediction; multiple potential futures can arise from a single current frame. To tackle this challenge, in this paper, we revisit the idea of stochastic video generation that learns to capture uncertainty in frame prediction and explore its effectiveness for representation learning. Specifically, we design a framework that trains a stochastic frame prediction model to learn temporal information between frames. Moreover, to learn dense information within each frame, we introduce an auxiliary masked image modeling objective along with a shared decoder architecture. We find this architecture allows for combining both objectives in a synergistic and compute-efficient manner. We demonstrate the effectiveness of our framework on a variety of tasks from video label propagation and vision-based robot learning domains, such as video segmentation, pose tracking, vision-based robotic locomotion, and manipulation tasks. Code is available on the project webpage: this https URL.

Comments:	International Conference on Machine Learning (ICML) 2024
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
Cite as:	arXiv:2406.07398 [cs.CV]
	(or arXiv:2406.07398v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2406.07398

Submission history

From: Younggyo Seo [view email]
[v1] Tue, 11 Jun 2024 16:05:15 UTC (1,864 KB)
[v2] Thu, 8 Aug 2024 19:48:10 UTC (1,864 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Visual Representation Learning with Stochastic Frame Prediction

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Visual Representation Learning with Stochastic Frame Prediction

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators