Can Large Language Models Reason? A Characterization via 3-SAT

Hazra, Rishi; Venturato, Gabriele; Martires, Pedro Zuidberg Dos; De Raedt, Luc

Computer Science > Artificial Intelligence

arXiv:2408.07215 (cs)

[Submitted on 13 Aug 2024 (v1), last revised 22 Oct 2024 (this version, v2)]

Title:Can Large Language Models Reason? A Characterization via 3-SAT

Authors:Rishi Hazra, Gabriele Venturato, Pedro Zuidberg Dos Martires, Luc De Raedt

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) have been touted as AI models possessing advanced reasoning abilities. However, recent works have shown that LLMs often bypass true reasoning using shortcuts, sparking skepticism. To study the reasoning capabilities in a principled fashion, we adopt a computational theory perspective and propose an experimental protocol centered on 3-SAT -- the prototypical NP-complete problem lying at the core of logical reasoning and constraint satisfaction tasks. Specifically, we examine the phase transitions in random 3-SAT and characterize the reasoning abilities of LLMs by varying the inherent hardness of the problem instances. Our experimental evidence shows that LLMs are incapable of performing true reasoning, as required for solving 3-SAT problems. Moreover, we observe significant performance variation based on the inherent hardness of the problems -- performing poorly on harder instances and vice versa. Importantly, we show that integrating external reasoners can considerably enhance LLM performance. By following a principled experimental protocol, our study draws concrete conclusions and moves beyond the anecdotal evidence often found in LLM reasoning research.

Subjects:	Artificial Intelligence (cs.AI)
Cite as:	arXiv:2408.07215 [cs.AI]
	(or arXiv:2408.07215v2 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2408.07215

Submission history

From: Rishi Hazra [view email]
[v1] Tue, 13 Aug 2024 21:54:10 UTC (14,565 KB)
[v2] Tue, 22 Oct 2024 21:44:03 UTC (24,540 KB)

Computer Science > Artificial Intelligence

Title:Can Large Language Models Reason? A Characterization via 3-SAT

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:Can Large Language Models Reason? A Characterization via 3-SAT

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators