Optimizing Relevance Maps of Vision Transformers Improves Robustness

Chefer, Hila; Schwartz, Idan; Wolf, Lior

Computer Science > Computer Vision and Pattern Recognition

arXiv:2206.01161 (cs)

[Submitted on 2 Jun 2022]

Title:Optimizing Relevance Maps of Vision Transformers Improves Robustness

Authors:Hila Chefer, Idan Schwartz, Lior Wolf

View PDF

Abstract:It has been observed that visual classification models often rely mostly on the image background, neglecting the foreground, which hurts their robustness to distribution changes. To alleviate this shortcoming, we propose to monitor the model's relevancy signal and manipulate it such that the model is focused on the foreground object. This is done as a finetuning step, involving relatively few samples consisting of pairs of images and their associated foreground masks. Specifically, we encourage the model's relevancy map (i) to assign lower relevance to background regions, (ii) to consider as much information as possible from the foreground, and (iii) we encourage the decisions to have high confidence. When applied to Vision Transformer (ViT) models, a marked improvement in robustness to domain shifts is observed. Moreover, the foreground masks can be obtained automatically, from a self-supervised variant of the ViT model itself; therefore no additional supervision is required.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2206.01161 [cs.CV]
	(or arXiv:2206.01161v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2206.01161

Submission history

From: Hila Chefer [view email]
[v1] Thu, 2 Jun 2022 17:24:48 UTC (37,602 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Optimizing Relevance Maps of Vision Transformers Improves Robustness

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Optimizing Relevance Maps of Vision Transformers Improves Robustness

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators