A Diffusion-based Method for Multi-turn Compositional Image Generation

Wang, Chao

Computer Science > Computer Vision and Pattern Recognition

arXiv:2304.02192 (cs)

[Submitted on 5 Apr 2023 (v1), last revised 14 Nov 2023 (this version, v2)]

Title:A Diffusion-based Method for Multi-turn Compositional Image Generation

Authors:Chao Wang

View PDF

Abstract:Multi-turn compositional image generation (M-CIG) is a challenging task that aims to iteratively manipulate a reference image given a modification text. While most of the existing methods for M-CIG are based on generative adversarial networks (GANs), recent advances in image generation have demonstrated the superiority of diffusion models over GANs. In this paper, we propose a diffusion-based method for M-CIG named conditional denoising diffusion with image compositional matching (CDD-ICM). We leverage CLIP as the backbone of image and text encoders, and incorporate a gated fusion mechanism, originally proposed for question answering, to compositionally fuse the reference image and the modification text at each turn of M-CIG. We introduce a conditioning scheme to generate the target image based on the fusion results. To prioritize the semantic quality of the generated target image, we learn an auxiliary image compositional match (ICM) objective, along with the conditional denoising diffusion (CDD) objective in a multi-task learning framework. Additionally, we also perform ICM guidance and classifier-free guidance to improve performance. Experimental results show that CDD-ICM achieves state-of-the-art results on two benchmark datasets for M-CIG, i.e., CoDraw and i-CLEVR.

Comments:	WACV 2024 3rd Workshop on Image/Video/Audio Quality in Computer Vision and Generative AI
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2304.02192 [cs.CV]
	(or arXiv:2304.02192v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2304.02192

Submission history

From: Chao Wang [view email]
[v1] Wed, 5 Apr 2023 02:13:42 UTC (2,809 KB)
[v2] Tue, 14 Nov 2023 02:01:38 UTC (2,810 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:A Diffusion-based Method for Multi-turn Compositional Image Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:A Diffusion-based Method for Multi-turn Compositional Image Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators