paper-with-me

홈 › Papers

Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models

2025-06-17 · Ling Li, Yao Zhou, Yuxuan Liang, Fugee Tsung, Jiaheng Wei

Previous methods for image geo-localization have typically treated the task as either classification or retrieval, often relying on black-box decisions that lack interpretability. The rise of large vision-language models (LVLMs) has enabled a rethinking of geo-localization as a reasoning-driven task grounded in visual cues. However, two major challenges persist. On the data side, existing reasoning-focused datasets are primarily based on street-view imagery, offering limited scene diversity and constrained viewpoints. On the modeling side, current approaches predominantly rely on supervised fine-tuning, which yields only marginal improvements in reasoning capabilities. To address these challenges, we propose a novel pipeline that constructs a reasoning-oriented geo-localization dataset, MP16-Reason, using diverse social media images. We introduce GLOBE, Group-relative policy optimization for Locatability assessment and Optimized visual-clue reasoning, yielding Bi-objective geo-Enhancement for the VLM in recognition and reasoning. GLOBE incorporates task-specific rewards that jointly enhance locatability assessment, visual clue reasoning, and geolocation accuracy. Both qualitative and quantitative results demonstrate that GLOBE outperforms state-of-the-art open-source LVLMs on geo-localization tasks, particularly in diverse visual scenes, while also generating more insightful and interpretable reasoning trajectories.

📄 PDF Abstract BibTeX arXiv:2506.14674

Code (0)

등록된 구현이 없습니다.

Tasks

geo-localization

Similar Papers 제목 키워드 기반

AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly Detection

2026-02-09 · Junru Zhang, Lang Feng, Haoran Shi, Xu Guo 외 arxiv

Time-series anomaly detection (TSAD) with multimodal large language models (MLLMs) is an emerging area, yet a persistent challenge remains: MLLMs rely on coarse time-series heuristics but struggle with multi-dimensional,…

Anomaly ClassificationReinforcement LearningAnomaly Detection

Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization

2026-04-17 · Siddhant Bharadwaj, Ashish Vashist, Fahimul Aleem, Shruti Vyas arxiv

Image geolocalization has traditionally been addressed through retrieval-based place recognition or geometry-based visual localization pipelines. Recent advances in Vision-Language Models (VLMs) have demonstrated strong …

Multimodal ReasoningVisual LocalizationImage Matching

Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention

2026-05-21 · Changyuan Tian, Zhicong Lu, Huaxing Liu, Xiang Wang 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for advancing complex reasoning in large language models, and recent work extends RLVR to multimodal large language models (MLLMs)…

Reinforcement LearningMultimodal Reasoning

Diagnosing Knowledge Conflict in Multimodal Long-Chain Reasoning

2026-02-16 · Jing Tang, Kun Wang, Haolang Lu, Hongjin Chen 외 arxiv

Multimodal large language models (MLLMs) in long chain-of-thought reasoning often fail when different knowledge sources provide conflicting signals. We formalize these failures under a unified notion of knowledge conflic…

Multimodal Reasoning

Reinforcing Generated Images via Meta-learning for One-Shot Fine-Grained Visual Recognition

2022-04-22 · Satoshi Tsutsui, Yanwei Fu, David Crandall

One-shot fine-grained visual recognition often suffers from the problem of having few training examples for new fine-grained classes. To alleviate this problem, off-the-shelf image generation techniques based on Generati…

DiversityFine-Grained Image ClassificationFine-Grained Visual Recognitionimage-classification+4