DATETIME: A new benchmark to measure LLM translation and reasoning capabilities

Gaere, Edward; Wangenheim, Florian

Computer Science > Neural and Evolutionary Computing

arXiv:2504.16155 (cs)

[Submitted on 22 Apr 2025]

Title:DATETIME: A new benchmark to measure LLM translation and reasoning capabilities

Authors:Edward Gaere, Florian Wangenheim

View PDF HTML (experimental)

Abstract:This paper introduces DATETIME, a new high-quality benchmark designed to evaluate the translation and reasoning abilities of a Large Language Model (LLM) on datetimes. A datetime is simply a date and a time, for example '11th.february.2023 ,1:12:31'. Datetimes are an interesting domain because they are intuitive and straightforward for humans to process but present significant challenges for LLMs. At the time of writing, no publicly available benchmark exists for systematically evaluating LLMs on datetime processing. Our experiments show that state-of-the-art models exhibit significant difficulty with tasks involving reasoning on datetimes, and that General Artificial Intelligence is still a distant aspiration. We hypothesize that working with datetimes necessitates translation and/or computation capabilities, and the tasks of the benchmark are organized accordingly. Significant dispersion in performance across models is observed with surprisingly poor performance even on apparently trivial tasks. Whilst frontier models such as ChatGPT, Claude and Llama3.1 have evidently been built and trained with datetime reasoning abilities, significant improvement is required for the open-source models.

Subjects:	Neural and Evolutionary Computing (cs.NE)
Cite as:	arXiv:2504.16155 [cs.NE]
	(or arXiv:2504.16155v1 [cs.NE] for this version)
	https://doi.org/10.48550/arXiv.2504.16155

Submission history

From: Edward Gaere [view email]
[v1] Tue, 22 Apr 2025 17:52:04 UTC (6,096 KB)

Computer Science > Neural and Evolutionary Computing

Title:DATETIME: A new benchmark to measure LLM translation and reasoning capabilities

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Neural and Evolutionary Computing

Title:DATETIME: A new benchmark to measure LLM translation and reasoning capabilities

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators