paper-with-me

홈 › Papers

Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models

2026-03-13 · Bo Yu, Fengze Yang, Yiming Liu, Chao Wang, Xuewen Luo, Taozhe Li, Ruimin Ke, Xiaofan Zhou, Chenxi Liu arxiv

The emergence of Vision-Language Models (VLMs) has introduced new paradigms for global image geo-localization through retrieval-augmented generation (RAG) and reasoning-driven inference. However, RAG methods are constrained by retrieval database quality, while reasoning-driven approaches fail to internalize image locatability, relying on inefficient, fixed-depth reasoning paths that increase hallucinations and degrade accuracy. To overcome these limitations, we introduce an Optimized Locatability Score that quantifies an image's suitability for deep reasoning in geo-localization. Using this metric, we curate Geo-ADAPT-51K, a locatability-stratified reasoning dataset enriched with augmented reasoning trajectories for complex visual scenes. Building on this foundation, we propose a two-stage Group Relative Policy Optimization (GRPO) curriculum with customized reward functions that regulate adaptive reasoning depth, visual grounding, and hierarchical geographical accuracy. Our framework, Geo-ADAPT, learns an adaptive reasoning policy, achieves state-of-the-art performance across multiple geo-localization benchmarks, and substantially reduces hallucinations by reasoning both adaptively and efficiently.

📄 PDF Abstract BibTeX arXiv:2603.13628

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models

2025-06-17 · Ling Li, Yao Zhou, Yuxuan Liang, Fugee Tsung 외

Previous methods for image geo-localization have typically treated the task as either classification or retrieval, often relying on black-box decisions that lack interpretability. The rise of large vision-language models…

geo-localization

Lifting Vision: Ground to Aerial Localization with Reasoning Guided Planning

2025-12-30 · Soham Pahari, M. Srinivas arxiv

Multimodal intelligence development recently show strong progress in visual understanding and high level reasoning. Though, most reasoning system still reply on textual information as the main medium for inference. This …

Contrastive LearningSpatial ReasoningVisual NavigationVisual Reasoning

Beyond Static Cropping: Layer-Adaptive Visual Localization and Decoding Enhancement

2026-02-04 · Zipeng Zhu, Zhanghao Hu, Qinglin Zhu, Yuxi Hong 외 arxiv

Large Vision-Language Models (LVLMs) have advanced rapidly by aligning visual patches with the text embedding space, but a fixed visual-token budget forces images to be resized to a uniform pretraining resolution, often …

Visual LocalizationQuestion AnsweringObject RecognitionVisual Grounding

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models

2026-02-26 · Yuxiang Shen, Hailong Huang, Zhenkun Gao, Xueheng Li 외 arxiv

Multimodal Large Language Models (MLLMs) are shifting towards "Thinking with Images" by actively exploring image details. While effective, large-scale training is computationally expensive, which has spurred growing inte…

Visual Reasoning

Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision

2026-03-28 · Yizhou Jin, Yuezhu Feng, Jinjin Zhang, Peng Wang 외 arxiv

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning and perceptual abilities for anomaly detection. However, most approaches remain confined to image-level anomaly detection and textu…

Reinforcement LearningAnomaly Detection