GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

August 07, 2026 Β· Grace Period Β· πŸ› CIKM2026

⏳ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno, Lynda Tamine arXiv ID 2608.07411 Category cs.AI: Artificial Intelligence Cross-listed cs.CL, cs.IR, cs.LG Citations 0 Venue CIKM2026
Abstract
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Artificial Intelligence