Interpreting Learned Feedback Patterns in Large Language Models

Marks, Luke; Abdullah, Amir; Neo, Clement; Arike, Rauno; Krueger, David; Torr, Philip; Barez, Fazl

Computer Science > Machine Learning

arXiv:2310.08164 (cs)

[Submitted on 12 Oct 2023 (v1), last revised 19 Aug 2024 (this version, v5)]

Title:Interpreting Learned Feedback Patterns in Large Language Models

Authors:Luke Marks, Amir Abdullah, Clement Neo, Rauno Arike, David Krueger, Philip Torr, Fazl Barez

View PDF HTML (experimental)

Abstract:Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferences in human feedback data. We coin the term \textit{Learned Feedback Pattern} (LFP) for patterns in an LLM's activations learned during RLHF that improve its performance on the fine-tuning task. We hypothesize that LLMs with LFPs accurately aligned to the fine-tuning feedback exhibit consistent activation patterns for outputs that would have received similar feedback during RLHF. To test this, we train probes to estimate the feedback signal implicit in the activations of a fine-tuned LLM. We then compare these estimates to the true feedback, measuring how accurate the LFPs are to the fine-tuning feedback. Our probes are trained on a condensed, sparse and interpretable representation of LLM activations, making it easier to correlate features of the input with our probe's predictions. We validate our probes by comparing the neural features they correlate with positive feedback inputs against the features GPT-4 describes and classifies as related to LFPs. Understanding LFPs can help minimize discrepancies between LLM behavior and training objectives, which is essential for the safety of LLMs.

Comments:	19 pages, 8 figures
Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2310.08164 [cs.LG]
	(or arXiv:2310.08164v5 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2310.08164

Submission history

From: Luke Marks [view email]
[v1] Thu, 12 Oct 2023 09:36:03 UTC (231 KB)
[v2] Tue, 28 Nov 2023 05:36:12 UTC (199 KB)
[v3] Mon, 5 Feb 2024 07:02:03 UTC (300 KB)
[v4] Wed, 7 Feb 2024 11:13:15 UTC (300 KB)
[v5] Mon, 19 Aug 2024 12:44:27 UTC (793 KB)

Computer Science > Machine Learning

Title:Interpreting Learned Feedback Patterns in Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Interpreting Learned Feedback Patterns in Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators