Self-Supervised Pre-training with Combined Datasets for 3D Perception in Autonomous Driving

Wang, Shumin; Yang, Zhuoran; Wang, Lidian; Tang, Zhipeng; Li, Heng; Pan, Lehan; Zhang, Sha; Peng, Jie; Ji, Jianmin; Zhang, Yanyong

Computer Science > Computer Vision and Pattern Recognition

arXiv:2504.12709 (cs)

[Submitted on 17 Apr 2025]

Title:Self-Supervised Pre-training with Combined Datasets for 3D Perception in Autonomous Driving

Authors:Shumin Wang, Zhuoran Yang, Lidian Wang, Zhipeng Tang, Heng Li, Lehan Pan, Sha Zhang, Jie Peng, Jianmin Ji, Yanyong Zhang

View PDF HTML (experimental)

Abstract:The significant achievements of pre-trained models leveraging large volumes of data in the field of NLP and 2D vision inspire us to explore the potential of extensive data pre-training for 3D perception in autonomous driving. Toward this goal, this paper proposes to utilize massive unlabeled data from heterogeneous datasets to pre-train 3D perception models. We introduce a self-supervised pre-training framework that learns effective 3D representations from scratch on unlabeled data, combined with a prompt adapter based domain adaptation strategy to reduce dataset bias. The approach significantly improves model performance on downstream tasks such as 3D object detection, BEV segmentation, 3D object tracking, and occupancy prediction, and shows steady performance increase as the training data volume scales up, demonstrating the potential of continually benefit 3D perception models for autonomous driving. We will release the source code to inspire further investigations in the community.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2504.12709 [cs.CV]
	(or arXiv:2504.12709v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2504.12709

Submission history

From: Zhuoran Yang [view email]
[v1] Thu, 17 Apr 2025 07:26:11 UTC (3,733 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Self-Supervised Pre-training with Combined Datasets for 3D Perception in Autonomous Driving

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Self-Supervised Pre-training with Combined Datasets for 3D Perception in Autonomous Driving

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators