paper-with-me

홈 › Papers

Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales

2025-10-13 · Zhaofang Qian, Hardy Chen, Zeyu Wang, Li Zhang, Zijun Wang, Xiaoke Huang, Hui Liu, Xianfeng Tang, Zeyu Zheng, Haoqin Tu, Cihang Xie, Yuyin Zhou arxiv

Vision-language models (VLMs) have advanced rapidly, yet their capacity for image-grounded geolocation in open-world conditions, a task that is challenging and of demand in real life, has not been comprehensively evaluated. We present EarthWhere, a comprehensive benchmark for VLM image geolocation that evaluates visual recognition, step-by-step reasoning, and evidence use. EarthWhere comprises 810 globally distributed images across two complementary geolocation scales: WhereCountry (i.e., 500 multiple-choice question-answering, with country-level answer and panoramas) and WhereStreet (i.e., 310 fine-grained street-level identification tasks requiring multi-step reasoning with optional web search). For evaluation, we adopt the final-prediction metrics: location accuracies within k km (Acc@k) for coordinates and hierarchical path scores for textual localization. Beyond this, we propose to explicitly score intermediate reasoning chains using human-verified key visual clues and a Shapley-reweighted thinking score that attributes credit to each clue's marginal contribution. We benchmark 13 state-of-the-art VLMs with web searching tools on our EarthWhere and report different types of final answer accuracies as well as the calibrated model thinking scores. Overall, Gemini-2.5-Pro achieves the best average accuracy at 56.32%, while the strongest open-weight model, GLM-4.5V, reaches 34.71%. We reveal that web search and reasoning do not guarantee improved performance when visual clues are limited, and models exhibit regional biases, achieving up to 42.7% higher scores in certain areas than others. These findings highlight not only the promise but also the persistent challenges of models to mitigate bias and achieve robust, fine-grained localization. We open-source our benchmark at https://github.com/UCSC-VLAA/EarthWhere.

📄 PDF Abstract BibTeX arXiv:2510.10880

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

2026-03-07 · Lance Legel, Qin Huang, Brandon Voelker, Daniel Neamati 외 arxiv

We present DeepEarth, a self-supervised multi-modal world model with Earth4D, a novel planetary-scale 4D space-time positional encoder. Earth4D extends 3D multi-resolution hash encoding to include time, efficiently scali…

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features

2025-01-17 · Enes Karanfil, Nevrez Imamoglu, Erkut Erdem, Aykut Erdem

Scene understanding in remote sensing often faces challenges in generating accurate representations for complex environments such as various land use areas or coastal regions, which may also include snow, clouds, or haze…

Language ModelingLanguage ModellingScene ClassificationScene Understanding

EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework

2026-01-06 · Junjue Wang, Yanfei Zhong, Zihang Chen, Zhuo Zheng 외 arxiv

Earth vision has achieved milestones in geospatial object recognition but lacks exploration in object-relational reasoning, limiting comprehensive scene understanding. To address this, a progressive Earth vision-language…

Visual Question AnsweringSemantic SegmentationRelational ReasoningScene Understanding

Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imagery

2026-02-26 · Minh Kha Do, Wei Xiang, Kang Han, Di Wu 외 arxiv

Vision-language foundation models (VLFMs) promise zero-shot and retrieval understanding for Earth observation. While operational satellite systems often lack full multi-spectral coverage, making RGB-only inference highly…

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

2026-08-13 · Yupan Ding, Jing Xiao, Zhenyuan Zhang, Chaofeng Chen 외 arxiv

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existin…

Spatial Reasoning