paper-with-me

Papers

Image-Based Geolocation Using Large Vision-Language Models

2024-08-18 · Yi Liu, Junchen Ding, Gelei Deng, Yuekang Li, Tianwei Zhang, Weisong Sun, Yaowen Zheng, Jingquan Ge, Yang Liu

Geolocation is now a vital aspect of modern life, offering numerous benefits but also presenting serious privacy concerns. The advent of large vision-language models (LVLMs) with advanced image-processing capabilities introduces new risks, as these models can inadvertently reveal sensitive geolocation information. This paper presents the first in-depth study analyzing the challenges posed by traditional deep learning and LVLM-based geolocation methods. Our findings reveal that LVLMs can accurately determine geolocations from images, even without explicit geographic training. To address these challenges, we introduce \tool{}, an innovative framework that significantly enhances image-based geolocation accuracy. \tool{} employs a systematic chain-of-thought (CoT) approach, mimicking human geoguessing strategies by carefully analyzing visual and contextual cues such as vehicle types, architectural styles, natural landscapes, and cultural elements. Extensive testing on a dataset of 50,000 ground-truth data points shows that \tool{} outperforms both traditional models and human benchmarks in accuracy. It achieves an impressive average score of 4550.5 in the GeoGuessr game, with an 85.37\% win rate, and delivers highly precise geolocation predictions, with the closest distances as accurate as 0.3 km. Furthermore, our study highlights issues related to dataset integrity, leading to the creation of a more robust dataset and a refined framework that leverages LVLMs' cognitive capabilities to improve geolocation precision. These findings underscore \tool{}'s superior ability to interpret complex visual data, the urgent need to address emerging security vulnerabilities posed by LVLMs, and the importance of responsible AI development to ensure user privacy protection.

📄 PDF Abstract BibTeX arXiv:2408.09474

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GaGA: Towards Interactive Global Geolocation Assistant

2024-12-12 · Zhiyang Dou, Zipeng Wang, Xumeng Han, Guorong Li 외

Global geolocation, which seeks to predict the geographical location of images captured anywhere in the world, is one of the most challenging tasks in the field of computer vision. In this paper, we introduce an innovati…

World Knowledge

GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents

2025-11-27 · Xinyu Zhang, Yixin Wu, Boyang Zhang, Chenhao Lin 외 arxiv

Images shared on social media often expose geographic cues. While early geolocation methods required expert effort and lacked generalization, the rise of Large Vision Language Models (LVLMs) now enables accurate geolocat…

Granular Privacy Control for Geolocation with Vision Language Models

2024-07-06 · Ethan Mendes, Yang Chen, James Hays, Sauvik Das 외

Vision Language Models (VLMs) are rapidly advancing in their capability to answer information-seeking questions. As these models are widely deployed in consumer applications, they could lead to new privacy risks due to e…

LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild

2024-05-30 · Zhiqiang Wang, Dejia Xu, Rana Muhammad Shahroz Khan, Yanbin Lin 외

Image geolocation is a critical task in various image-understanding applications. However, existing methods often fail when analyzing challenging, in-the-wild images. Inspired by the exceptional background knowledge of m…

Benchmarking

Evaluating Precise Geolocation Inference Capabilities of Vision Language Models

2025-02-20 · Neel Jay, Hieu Minh Nguyen, Trung Dung Hoang, Jacob Haimes

The prevalence of Vision-Language Models (VLMs) raises important questions about privacy in an era where visual information is increasingly available. While foundation VLMs demonstrate broad knowledge and learned capabil…