Pre-training Point Cloud Compact Model with Partial-aware Reconstruction

Zha, Yaohua; Wang, Yanzi; Dai, Tao; Xia, Shu-Tao

Computer Science > Computer Vision and Pattern Recognition

arXiv:2407.09344 (cs)

[Submitted on 12 Jul 2024]

Title:Pre-training Point Cloud Compact Model with Partial-aware Reconstruction

Authors:Yaohua Zha, Yanzi Wang, Tao Dai, Shu-Tao Xia

View PDF HTML (experimental)

Abstract:The pre-trained point cloud model based on Masked Point Modeling (MPM) has exhibited substantial improvements across various tasks. However, two drawbacks hinder their practical application. Firstly, the positional embedding of masked patches in the decoder results in the leakage of their central coordinates, leading to limited 3D representations. Secondly, the excessive model size of existing MPM methods results in higher demands for devices. To address these, we propose to pre-train Point cloud Compact Model with Partial-aware \textbf{R}econstruction, named Point-CPR. Specifically, in the decoder, we couple the vanilla masked tokens with their positional embeddings as randomly masked queries and introduce a partial-aware prediction module before each decoder layer to predict them from the unmasked partial. It prevents the decoder from creating a shortcut between the central coordinates of masked patches and their reconstructed coordinates, enhancing the robustness of models. We also devise a compact encoder composed of local aggregation and MLPs, reducing the parameters and computational requirements compared to existing Transformer-based encoders. Extensive experiments demonstrate that our model exhibits strong performance across various tasks, especially surpassing the leading MPM-based model PointGPT-B with only 2% of its parameters.

Comments:	arXiv admin note: text overlap with arXiv:2405.17149
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2407.09344 [cs.CV]
	(or arXiv:2407.09344v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2407.09344

Submission history

From: Yaohua Zha [view email]
[v1] Fri, 12 Jul 2024 15:18:14 UTC (2,177 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Pre-training Point Cloud Compact Model with Partial-aware Reconstruction

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Pre-training Point Cloud Compact Model with Partial-aware Reconstruction

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators