Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation

Englert, Brunó B.; Piva, Fabrizio J.; Kerssies, Tommie; de Geus, Daan; Dubbelman, Gijs

Computer Science > Computer Vision and Pattern Recognition

arXiv:2406.09896v1 (cs)

[Submitted on 14 Jun 2024 (this version), latest version 17 Jun 2024 (v2)]

Title:Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation

Authors:Brunó B. Englert, Fabrizio J. Piva, Tommie Kerssies, Daan de Geus, Gijs Dubbelman

View PDF HTML (experimental)

Abstract:Achieving robust generalization across diverse data domains remains a significant challenge in computer vision. This challenge is important in safety-critical applications, where deep-neural-network-based systems must perform reliably under various environmental conditions not seen during training. Our study investigates whether the generalization capabilities of Vision Foundation Models (VFMs) and Unsupervised Domain Adaptation (UDA) methods for the semantic segmentation task are complementary. Results show that combining VFMs with UDA has two main benefits: (a) it allows for better UDA performance while maintaining the out-of-distribution performance of VFMs, and (b) it makes certain time-consuming UDA components redundant, thus enabling significant inference speedups. Specifically, with equivalent model sizes, the resulting VFM-UDA method achieves an 8.4$\times$ speed increase over the prior non-VFM state of the art, while also improving performance by +1.2 mIoU in the UDA setting and by +6.1 mIoU in terms of out-of-distribution generalization. Moreover, when we use a VFM with 3.6$\times$ more parameters, the VFM-UDA approach maintains a 3.3$\times$ speed up, while improving the UDA performance by +3.1 mIoU and the out-of-distribution performance by +10.3 mIoU. These results underscore the significant benefits of combining VFMs with UDA, setting new standards and baselines for Unsupervised Domain Adaptation in semantic segmentation.

Comments:	CVPR 2024 Workshop Proceedings for the Second Workshop on Foundation Models
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2406.09896 [cs.CV]
	(or arXiv:2406.09896v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2406.09896

Submission history

From: Brunó B. Englert [view email]
[v1] Fri, 14 Jun 2024 10:13:37 UTC (1,405 KB)
[v2] Mon, 17 Jun 2024 09:51:40 UTC (1,405 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators