Stochastic positional embeddings improve masked image modeling

Bar, Amir; Bordes, Florian; Shocher, Assaf; Assran, Mahmoud; Vincent, Pascal; Ballas, Nicolas; Darrell, Trevor; Globerson, Amir; LeCun, Yann

Computer Science > Computer Vision and Pattern Recognition

arXiv:2308.00566 (cs)

[Submitted on 31 Jul 2023 (v1), last revised 27 Feb 2024 (this version, v2)]

Title:Stochastic positional embeddings improve masked image modeling

Authors:Amir Bar, Florian Bordes, Assaf Shocher, Mahmoud Assran, Pascal Vincent, Nicolas Ballas, Trevor Darrell, Amir Globerson, Yann LeCun

View PDF HTML (experimental)

Abstract:Masked Image Modeling (MIM) is a promising self-supervised learning approach that enables learning from unlabeled images. Despite its recent success, learning good representations through MIM remains challenging because it requires predicting the right semantic content in accurate locations. For example, given an incomplete picture of a dog, we can guess that there is a tail, but we cannot determine its exact location. In this work, we propose to incorporate location uncertainty into MIM by using stochastic positional embeddings (StoP). Specifically, we condition the model on stochastic masked token positions drawn from a Gaussian distribution. StoP reduces overfitting to location features and guides the model toward learning features that are more robust to location uncertainties. Quantitatively, StoP improves downstream MIM performance on a variety of downstream tasks, including $+1.7\%$ on ImageNet linear probing using ViT-B, and $+2.5\%$ for ViT-H using $1\%$ of the data.

Comments:	Code and models available in this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2308.00566 [cs.CV]
	(or arXiv:2308.00566v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2308.00566

Submission history

From: Amir Bar [view email]
[v1] Mon, 31 Jul 2023 17:59:08 UTC (10,636 KB)
[v2] Tue, 27 Feb 2024 18:59:14 UTC (2,527 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Stochastic positional embeddings improve masked image modeling

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Stochastic positional embeddings improve masked image modeling

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators