GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

Danish, Muhammad Sohail; Munir, Muhammad Akhtar; Shah, Syed Roshaan Ali; Kuckreja, Kartik; Khan, Fahad Shahbaz; Fraccaro, Paolo; Lacoste, Alexandre; Khan, Salman

Computer Science > Computer Vision and Pattern Recognition

arXiv:2411.19325 (cs)

[Submitted on 28 Nov 2024 (v1), last revised 12 Mar 2025 (this version, v2)]

Title:GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

Authors:Muhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Kartik Kuckreja, Fahad Shahbaz Khan, Paolo Fraccaro, Alexandre Lacoste, Salman Khan

View PDF HTML (experimental)

Abstract:While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications. Generic VLM benchmarks are not designed to handle the complexities of geospatial data, an essential component for applications such as environmental monitoring, urban planning, and disaster management. Key challenges in the geospatial domain include temporal change detection, large-scale object counting, tiny object detection, and understanding relationships between entities in remote sensing imagery. To bridge this gap, we present GEOBench-VLM, a comprehensive benchmark specifically designed to evaluate VLMs on geospatial tasks, including scene understanding, object counting, localization, fine-grained categorization, segmentation, and temporal analysis. Our benchmark features over 10,000 manually verified instructions and spanning diverse visual conditions, object types, and scales. We evaluate several state-of-the-art VLMs to assess performance on geospatial-specific challenges. The results indicate that although existing VLMs demonstrate potential, they face challenges when dealing with geospatial-specific tasks, highlighting the room for further improvements. Notably, the best-performing LLaVa-OneVision achieves only 41.7% accuracy on MCQs, slightly more than GPT-4o, which is approximately double the random guess performance. Our benchmark is publicly available at this https URL .

Comments:	This updated version includes revisions and additional analysis
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2411.19325 [cs.CV]
	(or arXiv:2411.19325v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2411.19325

Submission history

From: Muhammad Sohail Danish [view email]
[v1] Thu, 28 Nov 2024 18:59:56 UTC (14,902 KB)
[v2] Wed, 12 Mar 2025 19:28:05 UTC (13,171 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators