Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance

Lee, Hyunsoo; Kang, Minsoo; Han, Bohyung

Computer Science > Computer Vision and Pattern Recognition

arXiv:2412.15798 (cs)

[Submitted on 20 Dec 2024]

Title:Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance

Authors:Hyunsoo Lee, Minsoo Kang, Bohyung Han

View PDF

Abstract:We present a simple but effective training-free approach for text-driven image-to-image translation based on a pretrained text-to-image diffusion model. Our goal is to generate an image that aligns with the target task while preserving the structure and background of a source image. To this end, we derive the representation guidance with a combination of two objectives: maximizing the similarity to the target prompt based on the CLIP score and minimizing the structural distance to the source latent variable. This guidance improves the fidelity of the generated target image to the given target prompt while maintaining the structure integrity of the source image. To incorporate the representation guidance component, we optimize the target latent variable of diffusion model's reverse process with the guidance. Experimental results demonstrate that our method achieves outstanding image-to-image translation performance on various tasks when combined with the pretrained Stable Diffusion model.

Comments:	WACV 2025
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2412.15798 [cs.CV]
	(or arXiv:2412.15798v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2412.15798

Submission history

From: Hyunsoo Lee [view email]
[v1] Fri, 20 Dec 2024 11:15:31 UTC (37,279 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators