paper-with-me

Papers

GeoRC: A Benchmark for Geolocation Reasoning Chains

2026-01-29 · Mohit Talreja, Joshua Diao, Jim Thannikary James, Radu Casapu, Tejas Santanam, Ethan Mendes, Alan Ritter, Wei Xu, James Hays arxiv

Vision Language Models (VLMs) are good at recognizing the global location of a photograph -- their geolocation prediction accuracy rivals the best human experts. But many VLMs are startlingly bad at \textit{explaining} which image evidence led to their prediction, even when their location prediction is correct. In this paper, we introduce GeoRC, the first benchmark for geolocation reasoning chains sourced directly from Champion-tier GeoGuessr experts, including the reigning world champion. This benchmark consists of 800 ``ground truth'' reasoning chains across 500 query scenes from GeoGuessr maps, with expert chains addressing hundreds of different discriminative attributes, such as soil properties, architecture, and license plate shapes. We evaluate LLM-as-a-judge and VLM-as-a-judge strategies for scoring VLM-generated reasoning chains against our expert reasoning chains and find that Qwen 3 LLM-as-a-judge correlates best with human-expert scoring. Our benchmark reveals that while large, closed-source VLMs such as Gemini and GPT 5 rival human experts at predicting locations, they still lag behind human experts when it comes to producing auditable reasoning chains. Small open-weight VLMs such as Llama and Qwen catastrophically fail on our benchmark -- they perform only slightly better than a baseline in which an LLM hallucinates a reasoning chain with oracle knowledge of the photo location but \textit{no visual information at all}. We believe the gap between human experts and VLMs on this task points to VLM limitations at extracting fine-grained visual attributes from high resolution images. We open source our benchmark for the community to use.

📄 PDF Abstract BibTeX arXiv:2601.21278

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Skill-Conditioned Visual Geolocation for Vision-Language Models

2026-04-10 · Chenjie Yang, Yutian Jiang, Yutong Deng, Chenyu Wu arxiv

Vision-language models (VLMs) have shown a promising ability in image geolocation, but they still lack structured geographic reasoning and the capacity for autonomous self-evolution. Existing methods predominantly rely o…

Learning to Wander: Improving the Global Image Geolocation Ability of LMMs via Actionable Reasoning

2026-03-11 · Yushuo Zheng, Huiyu Duan, Zicheng Zhang, Xiaohong Liu 외 arxiv

Geolocation, the task of identifying the geographic location of an image, requires abundant world knowledge and complex reasoning abilities. Though advanced large multimodal models (LMMs) have shown superior aforemention…

Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales

2025-10-13 · Zhaofang Qian, Hardy Chen, Zeyu Wang, Li Zhang 외 arxiv

Vision-language models (VLMs) have advanced rapidly, yet their capacity for image-grounded geolocation in open-world conditions, a task that is challenging and of demand in real life, has not been comprehensively evaluat…

Simultaneous Optimization of Geodesics and Fréchet Means

2025-11-06 · Frederik Möbius Rygaard, Søren Hauberg, Steen Markvorsen arxiv

A central part of geometric statistics is to compute the Fréchet mean. This is a well-known intrinsic mean on a Riemannian manifold that minimizes the sum of squared Riemannian distances from the mean point to all other …

Interpretable Perception and Reasoning for Audiovisual Geolocation

2026-03-05 · Yiyang Su, Xiaoming Liu arxiv

While recent advances in Multimodal Large Language Models (MLLMs) have improved image-based localization, precise global geolocation remains a formidable challenge due to the inherent ambiguity of visual landscapes and t…

Image-Based LocalizationMultimodal Reasoning