Do Language Models Understand the Cognitive Tasks Given to Them? Investigations with the N-Back Paradigm

Hu, Xiaoyang; Lewis, Richard L.

Computer Science > Computation and Language

arXiv:2412.18120 (cs)

[Submitted on 24 Dec 2024 (v1), last revised 26 Dec 2024 (this version, v2)]

Title:Do Language Models Understand the Cognitive Tasks Given to Them? Investigations with the N-Back Paradigm

Authors:Xiaoyang Hu, Richard L. Lewis

View PDF HTML (experimental)

Abstract:Cognitive tasks originally developed for humans are now increasingly used to study language models. While applying these tasks is often straightforward, interpreting their results can be challenging. In particular, when a model underperforms, it is often unclear whether this results from a limitation in the cognitive ability being tested or a failure to understand the task itself. A recent study argues that GPT 3.5's declining performance on 2-back and 3-back tasks reflects a working memory capacity limit similar to humans (Gong et al., 2024). By analyzing a range of open-source language models of varying performance levels on these tasks, we show that the poor performance instead reflects a limitation in task comprehension and task set maintenance. In addition, we challenge the best-performing model with progressively harder versions of the task (up to 10-back) and experiment with alternative prompting strategies, before analyzing model attentions. Our larger aim is to contribute to the ongoing conversation around refining methodologies for the cognitive evaluation of language models.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2412.18120 [cs.CL]
	(or arXiv:2412.18120v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2412.18120

Submission history

From: Xiaoyang Hu [view email]
[v1] Tue, 24 Dec 2024 03:06:52 UTC (23,017 KB)
[v2] Thu, 26 Dec 2024 16:31:53 UTC (23,016 KB)

Computer Science > Computation and Language

Title:Do Language Models Understand the Cognitive Tasks Given to Them? Investigations with the N-Back Paradigm

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Do Language Models Understand the Cognitive Tasks Given to Them? Investigations with the N-Back Paradigm

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators