Benchmarking Systematic Relational Reasoning with Large Language and Reasoning Models

Khalid, Irtaza; Nourollah, Amir Masoud; Schockaert, Steven

Computer Science > Artificial Intelligence

arXiv:2503.23487 (cs)

[Submitted on 30 Mar 2025]

Title:Benchmarking Systematic Relational Reasoning with Large Language and Reasoning Models

Authors:Irtaza Khalid, Amir Masoud Nourollah, Steven Schockaert

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) have been found to struggle with systematic reasoning. Even on tasks where they appear to perform well, their performance often depends on shortcuts, rather than on genuine reasoning abilities, leading them to collapse on out-of-distribution examples. Post-training strategies based on reinforcement learning and chain-of-thought prompting have recently been hailed as a step change. However, little is still known about the potential of the resulting ``Large Reasoning Models'' (LRMs) beyond problem solving in mathematics and programming, where finding genuine out-of-distribution problems can be difficult. In this paper, we focus on tasks that require systematic reasoning about relational compositions, especially for qualitative spatial and temporal reasoning. These tasks allow us to control the difficulty of problem instances, and measure in a precise way to what extent models can generalise. We find that that the considered LLMs and LRMs overall perform poorly overall, albeit better than random chance.

Comments:	Submitted to ACL 2025
Subjects:	Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2503.23487 [cs.AI]
	(or arXiv:2503.23487v1 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2503.23487

Submission history

From: Irtaza Khalid [view email]
[v1] Sun, 30 Mar 2025 15:41:55 UTC (708 KB)

Computer Science > Artificial Intelligence

Title:Benchmarking Systematic Relational Reasoning with Large Language and Reasoning Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:Benchmarking Systematic Relational Reasoning with Large Language and Reasoning Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators