Pseudo Labelling for Enhanced Masked Autoencoders

Nandam, Srinivasa Rao; Atito, Sara; Feng, Zhenhua; Kittler, Josef; Awais, Muhammad

Computer Science > Computer Vision and Pattern Recognition

arXiv:2406.17450 (cs)

[Submitted on 25 Jun 2024]

Title:Pseudo Labelling for Enhanced Masked Autoencoders

Authors:Srinivasa Rao Nandam, Sara Atito, Zhenhua Feng, Josef Kittler, Muhammad Awais

View PDF HTML (experimental)

Abstract:Masked Image Modeling (MIM)-based models, such as SdAE, CAE, GreenMIM, and MixAE, have explored different strategies to enhance the performance of Masked Autoencoders (MAE) by modifying prediction, loss functions, or incorporating additional architectural components. In this paper, we propose an enhanced approach that boosts MAE performance by integrating pseudo labelling for both class and data tokens, alongside replacing the traditional pixel-level reconstruction with token-level reconstruction. This strategy uses cluster assignments as pseudo labels to promote instance-level discrimination within the network, while token reconstruction requires generation of discrete tokens encapturing local context. The targets for pseudo labelling and reconstruction needs to be generated by a teacher network. To disentangle the generation of target pseudo labels and the reconstruction of the token features, we decouple the teacher into two distinct models, where one serves as a labelling teacher and the other as a reconstruction teacher. This separation proves empirically superior to a single teacher, while having negligible impact on throughput and memory consumption. Incorporating pseudo-labelling as an auxiliary task has demonstrated notable improvements in ImageNet-1K and other downstream tasks, including classification, semantic segmentation, and detection.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2406.17450 [cs.CV]
	(or arXiv:2406.17450v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2406.17450

Submission history

From: Srinivasa Rao Nandam [view email]
[v1] Tue, 25 Jun 2024 10:41:45 UTC (9,202 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Pseudo Labelling for Enhanced Masked Autoencoders

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Pseudo Labelling for Enhanced Masked Autoencoders

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators