Gradient-based Jailbreak Images for Multimodal Fusion Models

Rando, Javier; Korevaar, Hannah; Brinkman, Erik; Evtimov, Ivan; Tramèr, Florian

Computer Science > Cryptography and Security

arXiv:2410.03489 (cs)

[Submitted on 4 Oct 2024 (v1), last revised 23 Oct 2024 (this version, v2)]

Title:Gradient-based Jailbreak Images for Multimodal Fusion Models

Authors:Javier Rando, Hannah Korevaar, Erik Brinkman, Ivan Evtimov, Florian Tramèr

View PDF HTML (experimental)

Abstract:Augmenting language models with image inputs may enable more effective jailbreak attacks through continuous optimization, unlike text inputs that require discrete optimization. However, new multimodal fusion models tokenize all input modalities using non-differentiable functions, which hinders straightforward attacks. In this work, we introduce the notion of a tokenizer shortcut that approximates tokenization with a continuous function and enables continuous optimization. We use tokenizer shortcuts to create the first end-to-end gradient image attacks against multimodal fusion models. We evaluate our attacks on Chameleon models and obtain jailbreak images that elicit harmful information for 72.5% of prompts. Jailbreak images outperform text jailbreaks optimized with the same objective and require 3x lower compute budget to optimize 50x more input tokens. Finally, we find that representation engineering defenses, like Circuit Breakers, trained only on text attacks can effectively transfer to adversarial image inputs.

Subjects:	Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2410.03489 [cs.CR]
	(or arXiv:2410.03489v2 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2410.03489

Submission history

From: Javier Rando [view email]
[v1] Fri, 4 Oct 2024 14:59:39 UTC (1,018 KB)
[v2] Wed, 23 Oct 2024 13:38:29 UTC (1,018 KB)

Computer Science > Cryptography and Security

Title:Gradient-based Jailbreak Images for Multimodal Fusion Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:Gradient-based Jailbreak Images for Multimodal Fusion Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators