Measuring Faithfulness of Chains of Thought by Unlearning Reasoning Steps

Tutek, Martin; Chaleshtori, Fateme Hashemi; Marasović, Ana; Belinkov, Yonatan

Computer Science > Computation and Language

arXiv:2502.14829 (cs)

[Submitted on 20 Feb 2025]

Title:Measuring Faithfulness of Chains of Thought by Unlearning Reasoning Steps

Authors:Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasović, Yonatan Belinkov

View PDF HTML (experimental)

Abstract:When prompted to think step-by-step, language models (LMs) produce a chain of thought (CoT), a sequence of reasoning steps that the model supposedly used to produce its prediction. However, despite much work on CoT prompting, it is unclear if CoT reasoning is faithful to the models' parameteric beliefs. We introduce a framework for measuring parametric faithfulness of generated reasoning, and propose Faithfulness by Unlearning Reasoning steps (FUR), an instance of this framework. FUR erases information contained in reasoning steps from model parameters. We perform experiments unlearning CoTs of four LMs prompted on four multi-choice question answering (MCQA) datasets. Our experiments show that FUR is frequently able to change the underlying models' prediction by unlearning key steps, indicating when a CoT is parametrically faithful. Further analysis shows that CoTs generated by models post-unlearning support different answers, hinting at a deeper effect of unlearning. Importantly, CoT steps identified as important by FUR do not align well with human notions of plausbility, emphasizing the need for specialized alignment

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2502.14829 [cs.CL]
	(or arXiv:2502.14829v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2502.14829

Submission history

From: Martin Tutek [view email]
[v1] Thu, 20 Feb 2025 18:45:05 UTC (920 KB)

Computer Science > Computation and Language

Title:Measuring Faithfulness of Chains of Thought by Unlearning Reasoning Steps

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Measuring Faithfulness of Chains of Thought by Unlearning Reasoning Steps

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators