Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate

Qi, Senmao; Zou, Yifei; Li, Peng; Lin, Ziyi; Cheng, Xiuzhen; Yu, Dongxiao

Computer Science > Cryptography and Security

arXiv:2504.16489 (cs)

[Submitted on 23 Apr 2025]

Title:Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate

Authors:Senmao Qi, Yifei Zou, Peng Li, Ziyi Lin, Xiuzhen Cheng, Dongxiao Yu

View PDF

Abstract:Multi-Agent Debate (MAD), leveraging collaborative interactions among Large Language Models (LLMs), aim to enhance reasoning capabilities in complex tasks. However, the security implications of their iterative dialogues and role-playing characteristics, particularly susceptibility to jailbreak attacks eliciting harmful content, remain critically underexplored. This paper systematically investigates the jailbreak vulnerabilities of four prominent MAD frameworks built upon leading commercial LLMs (GPT-4o, GPT-4, GPT-3.5-turbo, and DeepSeek) without compromising internal agents. We introduce a novel structured prompt-rewriting framework specifically designed to exploit MAD dynamics via narrative encapsulation, role-driven escalation, iterative refinement, and rhetorical obfuscation. Our extensive experiments demonstrate that MAD systems are inherently more vulnerable than single-agent setups. Crucially, our proposed attack methodology significantly amplifies this fragility, increasing average harmfulness from 28.14% to 80.34% and achieving attack success rates as high as 80% in certain scenarios. These findings reveal intrinsic vulnerabilities in MAD architectures and underscore the urgent need for robust, specialized defenses prior to real-world deployment.

Comments:	33 pages, 5 figures
Subjects:	Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2504.16489 [cs.CR]
	(or arXiv:2504.16489v1 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2504.16489

Submission history

From: Senmao Qi [view email]
[v1] Wed, 23 Apr 2025 08:01:50 UTC (3,830 KB)

Computer Science > Cryptography and Security

Title:Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators