Unleashing Vecset Diffusion Model for Fast Shape Generation

Lai, Zeqiang; Zhao, Yunfei; Zhao, Zibo; Liu, Haolin; Wang, Fuyun; Shi, Huiwen; Yang, Xianghui; Lin, Qingxiang; Huang, Jingwei; Liu, Yuhong; Jiang, Jie; Guo, Chunchao; Yue, Xiangyu

Computer Science > Computer Vision and Pattern Recognition

arXiv:2503.16302 (cs)

[Submitted on 20 Mar 2025 (v1), last revised 26 Mar 2025 (this version, v2)]

Title:Unleashing Vecset Diffusion Model for Fast Shape Generation

Authors:Zeqiang Lai, Yunfei Zhao, Zibo Zhao, Haolin Liu, Fuyun Wang, Huiwen Shi, Xianghui Yang, Qingxiang Lin, Jingwei Huang, Yuhong Liu, Jie Jiang, Chunchao Guo, Xiangyu Yue

View PDF HTML (experimental)

Abstract:3D shape generation has greatly flourished through the development of so-called "native" 3D diffusion, particularly through the Vecset Diffusion Model (VDM). While recent advancements have shown promising results in generating high-resolution 3D shapes, VDM still struggles with high-speed generation. Challenges exist because of difficulties not only in accelerating diffusion sampling but also VAE decoding in VDM, areas under-explored in previous works. To address these challenges, we present FlashVDM, a systematic framework for accelerating both VAE and DiT in VDM. For DiT, FlashVDM enables flexible diffusion sampling with as few as 5 inference steps and comparable quality, which is made possible by stabilizing consistency distillation with our newly introduced Progressive Flow Distillation. For VAE, we introduce a lightning vecset decoder equipped with Adaptive KV Selection, Hierarchical Volume Decoding, and Efficient Network Design. By exploiting the locality of the vecset and the sparsity of shape surface in the volume, our decoder drastically lowers FLOPs, minimizing the overall decoding overhead. We apply FlashVDM to Hunyuan3D-2 to obtain Hunyuan3D-2 Turbo. Through systematic evaluation, we show that our model significantly outperforms existing fast 3D generation methods, achieving comparable performance to the state-of-the-art while reducing inference time by over 45x for reconstruction and 32x for generation. Code and models are available at this https URL.

Comments:	Technical report
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV)
Cite as:	arXiv:2503.16302 [cs.CV]
	(or arXiv:2503.16302v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2503.16302

Submission history

From: Zeqiang Lai [view email]
[v1] Thu, 20 Mar 2025 16:23:44 UTC (25,994 KB)
[v2] Wed, 26 Mar 2025 15:08:12 UTC (25,994 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Unleashing Vecset Diffusion Model for Fast Shape Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Unleashing Vecset Diffusion Model for Fast Shape Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators