STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes

Yang, Jiawei; Huang, Jiahui; Chen, Yuxiao; Wang, Yan; Li, Boyi; You, Yurong; Sharma, Apoorva; Igl, Maximilian; Karkus, Peter; Xu, Danfei; Ivanovic, Boris; Wang, Yue; Pavone, Marco

Computer Science > Computer Vision and Pattern Recognition

arXiv:2501.00602 (cs)

[Submitted on 31 Dec 2024]

Title:STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes

Authors:Jiawei Yang, Jiahui Huang, Yuxiao Chen, Yan Wang, Boyi Li, Yurong You, Apoorva Sharma, Maximilian Igl, Peter Karkus, Danfei Xu, Boris Ivanovic, Yue Wang, Marco Pavone

View PDF HTML (experimental)

Abstract:We present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often rely on per-scene optimization, dense observations across space and time, and strong motion supervision, resulting in lengthy optimization times, limited generalization to novel views or scenes, and degenerated quality caused by noisy pseudo-labels for dynamics. To address these challenges, STORM leverages a data-driven Transformer architecture that directly infers dynamic 3D scene representations--parameterized by 3D Gaussians and their velocities--in a single forward pass. Our key design is to aggregate 3D Gaussians from all frames using self-supervised scene flows, transforming them to the target timestep to enable complete (i.e., "amodal") reconstructions from arbitrary viewpoints at any moment in time. As an emergent property, STORM automatically captures dynamic instances and generates high-quality masks using only reconstruction losses. Extensive experiments on public datasets show that STORM achieves precise dynamic scene reconstruction, surpassing state-of-the-art per-scene optimization methods (+4.3 to 6.6 PSNR) and existing feed-forward approaches (+2.1 to 4.7 PSNR) in dynamic regions. STORM reconstructs large-scale outdoor scenes in 200ms, supports real-time rendering, and outperforms competitors in scene flow estimation, improving 3D EPE by 0.422m and Acc5 by 28.02%. Beyond reconstruction, we showcase four additional applications of our model, illustrating the potential of self-supervised learning for broader dynamic scene understanding.

Comments:	Project page at: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as:	arXiv:2501.00602 [cs.CV]
	(or arXiv:2501.00602v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2501.00602

Submission history

From: Jiawei Yang [view email]
[v1] Tue, 31 Dec 2024 18:59:58 UTC (11,549 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators