Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

Ma, Jiachen; Cao, Anda; Xiao, Zhiqing; Zhang, Jie; Ye, Chao; Zhao, Junbo

Computer Science > Cryptography and Security

arXiv:2404.02928v2 (cs)

[Submitted on 2 Apr 2024 (v1), revised 2 Jun 2024 (this version, v2), latest version 4 Sep 2024 (v3)]

Title:Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

Authors:Jiachen Ma, Anda Cao, Zhiqing Xiao, Jie Zhang, Chao Ye, Junbo Zhao

View PDF HTML (experimental)

Abstract:Text-to-Image (T2I) models have received widespread attention due to their remarkable generation capabilities. However, concerns have been raised about the ethical implications of the models in generating Not Safe for Work (NSFW) images because NSFW images may cause discomfort to people or be used for illegal purposes. To mitigate the generation of such images, T2I models deploy various types of safety checkers. However, they still cannot completely prevent the generation of NSFW images. In this paper, we propose the Jailbreak Prompt Attack (JPA) - an automatic attack framework. We aim to maintain prompts that bypass safety checkers while preserving the semantics of the original images. Specifically, we aim to find prompts that can bypass safety checkers because of the robustness of the text space. Our evaluation demonstrates that JPA successfully bypasses both online services with closed-box safety checkers and offline defenses safety checkers to generate NSFW images.

Subjects:	Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2404.02928 [cs.CR]
	(or arXiv:2404.02928v2 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2404.02928

Submission history

From: Jiachen Ma [view email]
[v1] Tue, 2 Apr 2024 09:49:35 UTC (20,018 KB)
[v2] Sun, 2 Jun 2024 12:36:39 UTC (8,381 KB)
[v3] Wed, 4 Sep 2024 06:40:12 UTC (26,394 KB)

Computer Science > Cryptography and Security

Title:Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators