paper-with-me

Papers

GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization

2025-11-19 · Yikun Wang, Zuyan Liu, Ziyi Wang, Han Hu, Pengfei Liu, Yongming Rao arxiv

Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purpose agentic models. In this work, we revisit the geolocalization task, which requires not only nuanced visual grounding but also web search to confirm or refine hypotheses during reasoning. Since existing geolocalization benchmarks fail to meet the need for high-resolution imagery and the localization challenge for deep agentic reasoning, we curate GeoBench, a benchmark that includes photos and panoramas from around the world, along with a subset of satellite images of different cities to rigorously evaluate the geolocalization ability of agentic models. We also propose GeoVista, an agentic model that seamlessly integrates tool invocation within the reasoning loop, including an image-zoom-in tool to magnify regions of interest and a web-search tool to retrieve related web information. We develop a complete training pipeline for it, including a cold-start supervised fine-tuning (SFT) stage to learn reasoning patterns and tool-use priors, followed by a reinforcement learning (RL) stage to further enhance reasoning ability. We adopt a hierarchical reward to leverage multi-level geographical information and improve overall geolocalization performance. Experimental results show that GeoVista surpasses other open-source agentic models on the geolocalization task greatly and achieves performance comparable to closed-source models such as Gemini-2.5-flash and GPT-5 on most metrics.

📄 PDF Abstract BibTeX arXiv:2511.15705

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningImage ManipulationVisual GroundingVisual Reasoning

Similar Papers 제목 키워드 기반

Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization

2026-01-08 · Yuxiang Ji, Yong Wang, Ziyu Ma, Yiming Hu 외 arxiv

The image geolocalization task aims to predict the location where an image was taken anywhere on Earth using visual clues. Existing large vision-language model (LVLM) approaches leverage world knowledge, chain-of-thought…

Reinforcement Learning

GeoVista: Visually Grounded Active Perception for Vision-Language Understanding of Ultra-High-Resolution Remote Sensing Images

2026-05-14 · Jiashun Zhu, Ronghao Fu, Jiasen Hu, Jing Huang 외 arxiv

Interpreting ultra-high-resolution (UHR) remote sensing images requires models to search for sparse and tiny visual evidence across large-scale scenes. Existing remote sensing vision-language models can inspect local reg…

GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation

2026-05-14 · Yuhao Liu, Sadeer Al-Kindi, Ashok Veeraraghavan, Guha Balakrishnan arxiv

Large-scale pretraining on Earth observation imagery has yielded powerful representations of the natural and built environment. However, most existing geospatial foundation models do not directly model the structured soc…

Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization

2026-04-17 · Siddhant Bharadwaj, Ashish Vashist, Fahimul Aleem, Shruti Vyas arxiv

Image geolocalization has traditionally been addressed through retrieval-based place recognition or geometry-based visual localization pipelines. Recent advances in Vision-Language Models (VLMs) have demonstrated strong …

Multimodal ReasoningVisual LocalizationImage Matching

GeoSearch: Augmenting Worldwide Geolocalization with Web-Scale Reverse Image Search and Image Matching

2026-04-28 · Tung-Duong Le-Duc, Hoang-Quoc Nguyen-Son, Minh-Son Dao arxiv

Worldwide image geolocalization, which aims to predict the GPS coordinates of any image on Earth, remains challenging due to global visual diversity. Recent generative approaches based on Retrieval-Augmented Generation (…

Image Matching