MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability

Liu, Buyu; Wang, Kai; Liu, Yansong; Bao, Jun; Han, Tingting; Yu, Jun

Computer Science > Computer Vision and Pattern Recognition

arXiv:2407.19468 (cs)

[Submitted on 28 Jul 2024]

Title:MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability

Authors:Buyu Liu, Kai Wang, Yansong Liu, Jun Bao, Tingting Han, Jun Yu

View PDF HTML (experimental)

Abstract:This work aims to address the multi-view perspective RGB generation from text prompts given Bird-Eye-View(BEV) semantics. Unlike prior methods that neglect layout consistency, lack the ability to handle detailed text prompts, or are incapable of generalizing to unseen view points, MVPbev simultaneously generates cross-view consistent images of different perspective views with a two-stage design, allowing object-level control and novel view generation at test-time. Specifically, MVPbev firstly projects given BEV semantics to perspective view with camera parameters, empowering the model to generalize to unseen view points. Then we introduce a multi-view attention module where special initialization and de-noising processes are introduced to explicitly enforce local consistency among overlapping views w.r.t. cross-view homography. Last but not least, MVPbev further allows test-time instance-level controllability by refining a pre-trained text-to-image diffusion model. Our extensive experiments on NuScenes demonstrate that our method is capable of generating high-resolution photorealistic images from text descriptions with thousands of training samples, surpassing the state-of-the-art methods under various evaluation metrics. We further demonstrate the advances of our method in terms of generalizability and controllability with the help of novel evaluation metrics and comprehensive human analysis. Our code, data, and model can be found in \url{this https URL}.

Comments:	Accepted by ACM MM24
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
Cite as:	arXiv:2407.19468 [cs.CV]
	(or arXiv:2407.19468v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2407.19468

Submission history

From: Buyu Liu [view email]
[v1] Sun, 28 Jul 2024 11:39:40 UTC (35,561 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators