What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Jain, Samyak; Lubana, Ekdeep Singh; Oksuz, Kemal; Joy, Tom; Torr, Philip H. S.; Sanyal, Amartya; Dokania, Puneet K.

Computer Science > Machine Learning

arXiv:2407.10264 (cs)

[Submitted on 14 Jul 2024 (v1), last revised 21 Aug 2024 (this version, v3)]

Title:What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Authors:Samyak Jain, Ekdeep Singh Lubana, Kemal Oksuz, Tom Joy, Philip H.S. Torr, Amartya Sanyal, Puneet K. Dokania

View PDF

Abstract:Safety fine-tuning helps align Large Language Models (LLMs) with human preferences for their safe deployment. To better understand the underlying factors that make models safe via safety fine-tuning, we design a synthetic data generation framework that captures salient aspects of an unsafe input by modeling the interaction between the task the model is asked to perform (e.g., "design") versus the specific concepts the task is asked to be performed upon (e.g., a "cycle" vs. a "bomb"). Using this, we investigate three well-known safety fine-tuning methods -- supervised safety fine-tuning, direct preference optimization, and unlearning -- and provide significant evidence demonstrating that these methods minimally transform MLP weights to specifically align unsafe inputs into its weights' null space. This yields a clustering of inputs based on whether the model deems them safe or not. Correspondingly, when an adversarial input (e.g., a jailbreak) is provided, its activations are closer to safer samples, leading to the model processing such an input as if it were safe. We validate our findings, wherever possible, on real-world models -- specifically, Llama-2 7B and Llama-3 8B.

Comments:	Preprint
Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as:	arXiv:2407.10264 [cs.LG]
	(or arXiv:2407.10264v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2407.10264

Submission history

From: Ekdeep Singh Lubana [view email]
[v1] Sun, 14 Jul 2024 16:12:57 UTC (9,343 KB)
[v2] Tue, 16 Jul 2024 19:19:15 UTC (9,343 KB)
[v3] Wed, 21 Aug 2024 15:12:37 UTC (9,343 KB)

Computer Science > Machine Learning

Title:What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators