Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion

Liang, Li; Akhtar, Naveed; Vice, Jordan; Kong, Xiangrui; Mian, Ajmal Saeed

Computer Science > Computer Vision and Pattern Recognition

arXiv:2501.07260 (cs)

[Submitted on 13 Jan 2025]

Title:Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion

Authors:Li Liang, Naveed Akhtar, Jordan Vice, Xiangrui Kong, Ajmal Saeed Mian

View PDF HTML (experimental)

Abstract:3D semantic scene completion is critical for multiple downstream tasks in autonomous systems. It estimates missing geometric and semantic information in the acquired scene data. Due to the challenging real-world conditions, this task usually demands complex models that process multi-modal data to achieve acceptable performance. We propose a unique neural model, leveraging advances from the state space and diffusion generative modeling to achieve remarkable 3D semantic scene completion performance with monocular image input. Our technique processes the data in the conditioned latent space of a variational autoencoder where diffusion modeling is carried out with an innovative state space technique. A key component of our neural network is the proposed Skimba (Skip Mamba) denoiser, which is adept at efficiently processing long-sequence data. The Skimba diffusion model is integral to our 3D scene completion network, incorporating a triple Mamba structure, dimensional decomposition residuals and varying dilations along three directions. We also adopt a variant of this network for the subsequent semantic segmentation stage of our method. Extensive evaluation on the standard SemanticKITTI and SSCBench-KITTI360 datasets show that our approach not only outperforms other monocular techniques by a large margin, it also achieves competitive performance against stereo methods. The code is available at this https URL

Comments:	Accepted by AAAI 2025
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2501.07260 [cs.CV]
	(or arXiv:2501.07260v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2501.07260

Submission history

From: Xiangrui Kong Ray [view email]
[v1] Mon, 13 Jan 2025 12:18:58 UTC (966 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators