Back to the Color: Learning Depth to Specific Color Transformation for Unsupervised Depth Estimation

Zhu, Yufan; Ran, Chongzhi; Feng, Mingtao; Wu, Fangfang; Dong, Le; Dong, Weisheng; López, Antonio M.; Shi, Guangming

Computer Science > Computer Vision and Pattern Recognition

arXiv:2406.07741v3 (cs)

[Submitted on 11 Jun 2024 (v1), revised 3 Jul 2024 (this version, v3), latest version 15 Oct 2024 (v6)]

Title:Back to the Color: Learning Depth to Specific Color Transformation for Unsupervised Depth Estimation

Authors:Yufan Zhu, Chongzhi Ran, Mingtao Feng, Fangfang Wu, Le Dong, Weisheng Dong, Antonio M. López, Guangming Shi

View PDF HTML (experimental)

Abstract:Virtual engines can generate dense depth maps for various synthetic scenes, making them invaluable for training depth estimation models. However, discrepancies between synthetic and real-world colors pose significant challenges for depth estimation in real-world scenes, especially in complex and uncertain environments encountered in unsupervised monocular depth estimation tasks. To address this issue, we propose Back2Color, a framework that predicts realistic colors from depth using a model trained on real-world data, thus transforming synthetic colors into their real-world counterparts. Additionally, we introduce the Syn-Real CutMix method for joint training with both real-world unsupervised and synthetic supervised depth samples, enhancing monocular depth estimation performance in real-world scenes. Furthermore, to mitigate the impact of non-rigid motions on depth estimation, we present an auto-learning uncertainty temporal-spatial fusion method (Auto-UTSF), which leverages the strengths of unsupervised learning in both temporal and spatial dimensions. We also designed VADepth, based on the Vision Attention Network, which offers lower computational complexity and higher accuracy than transformers. Our Back2Color framework achieves state-of-the-art performance on the Kitti dataset, as evidenced by improvements in performance metrics and the production of fine-grained details. This is particularly evident on more challenging datasets such as Cityscapes for unsupervised depth estimation.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2406.07741 [cs.CV]
	(or arXiv:2406.07741v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2406.07741

Submission history

From: Yufan Zhu [view email]
[v1] Tue, 11 Jun 2024 21:55:20 UTC (41,293 KB)
[v2] Sun, 23 Jun 2024 21:54:26 UTC (39,841 KB)
[v3] Wed, 3 Jul 2024 14:08:22 UTC (41,031 KB)
[v4] Fri, 26 Jul 2024 08:48:47 UTC (41,032 KB)
[v5] Tue, 8 Oct 2024 07:43:54 UTC (41,032 KB)
[v6] Tue, 15 Oct 2024 07:27:28 UTC (41,030 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Back to the Color: Learning Depth to Specific Color Transformation for Unsupervised Depth Estimation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Back to the Color: Learning Depth to Specific Color Transformation for Unsupervised Depth Estimation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators