MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models

Ding, Peng; Fang, Jiading; Li, Peng; Wang, Kangrui; Zhou, Xiaochen; Yu, Mo; Li, Jing; Walter, Matthew R.; Mei, Hongyuan

Computer Science > Computation and Language

arXiv:2403.19913 (cs)

[Submitted on 29 Mar 2024 (v1), last revised 8 Aug 2024 (this version, v2)]

Title:MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models

Authors:Peng Ding, Jiading Fang, Peng Li, Kangrui Wang, Xiaochen Zhou, Mo Yu, Jing Li, Matthew R. Walter, Hongyuan Mei

View PDF HTML (experimental)

Abstract:Large language models such as ChatGPT and GPT-4 have recently achieved astonishing performance on a variety of natural language processing tasks. In this paper, we propose MANGO, a benchmark to evaluate their capabilities to perform text-based mapping and navigation. Our benchmark includes 53 mazes taken from a suite of textgames: each maze is paired with a walkthrough that visits every location but does not cover all possible paths. The task is question-answering: for each maze, a large language model reads the walkthrough and answers hundreds of mapping and navigation questions such as "How should you go to Attic from West of House?" and "Where are we if we go north and east from Cellar?". Although these questions are easy to humans, it turns out that even GPT-4, the best-to-date language model, performs poorly at answering them. Further, our experiments suggest that a strong mapping and navigation ability would benefit large language models in performing relevant downstream tasks, such as playing textgames. Our MANGO benchmark will facilitate future research on methods that improve the mapping and navigation capabilities of language models. We host our leaderboard, data, code, and evaluation program at this https URL and this https URL.

Comments:	COLM 2024 camera-ready
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
Cite as:	arXiv:2403.19913 [cs.CL]
	(or arXiv:2403.19913v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2403.19913

Submission history

From: Kangrui Wang [view email]
[v1] Fri, 29 Mar 2024 01:53:24 UTC (8,281 KB)
[v2] Thu, 8 Aug 2024 06:38:31 UTC (8,305 KB)

Computer Science > Computation and Language

Title:MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators