An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training

Zhang, Haiming; Xue, Ying; Yan, Xu; Zhang, Jiacheng; Qiu, Weichao; Bai, Dongfeng; Liu, Bingbing; Cui, Shuguang; Li, Zhen

Computer Science > Computer Vision and Pattern Recognition

arXiv:2412.13772 (cs)

[Submitted on 18 Dec 2024]

Title:An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training

Authors:Haiming Zhang, Ying Xue, Xu Yan, Jiacheng Zhang, Weichao Qiu, Dongfeng Bai, Bingbing Liu, Shuguang Cui, Zhen Li

View PDF HTML (experimental)

Abstract:The field of autonomous driving is experiencing a surge of interest in world models, which aim to predict potential future scenarios based on historical observations. In this paper, we introduce DFIT-OccWorld, an efficient 3D occupancy world model that leverages decoupled dynamic flow and image-assisted training strategy, substantially improving 4D scene forecasting performance. To simplify the training process, we discard the previous two-stage training strategy and innovatively reformulate the occupancy forecasting problem as a decoupled voxels warping process. Our model forecasts future dynamic voxels by warping existing observations using voxel flow, whereas static voxels are easily obtained through pose transformation. Moreover, our method incorporates an image-assisted training paradigm to enhance prediction reliability. Specifically, differentiable volume rendering is adopted to generate rendered depth maps through predicted future volumes, which are adopted in render-based photometric consistency. Experiments demonstrate the effectiveness of our approach, showcasing its state-of-the-art performance on the nuScenes and OpenScene benchmarks for 4D occupancy forecasting, end-to-end motion planning and point cloud forecasting. Concretely, it achieves state-of-the-art performances compared to existing 3D world models while incurring substantially lower computational costs.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2412.13772 [cs.CV]
	(or arXiv:2412.13772v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2412.13772

Submission history

From: Haiming Zhang [view email]
[v1] Wed, 18 Dec 2024 12:10:33 UTC (6,162 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators