paper-with-me

홈 › Papers

Assessing the Geolocation Capabilities, Limitations and Societal Risks of Generative Vision-Language Models

2025-08-27 · Oliver Grainge, Sania Waheed, Jack Stilgoe, Michael Milford, Shoaib Ehsan arxiv

Geo-localization is the task of identifying the location of an image using visual cues alone. It has beneficial applications, such as improving disaster response, enhancing navigation, and geography education. Recently, Vision-Language Models (VLMs) are increasingly demonstrating capabilities as accurate image geo-locators. This brings significant privacy risks, including those related to stalking and surveillance, considering the widespread uses of AI models and sharing of photos on social media. The precision of these models is likely to improve in the future. Despite these risks, there is little work on systematically evaluating the geolocation precision of Generative VLMs, their limits and potential for unintended inferences. To bridge this gap, we conduct a comprehensive assessment of the geolocation capabilities of 25 state-of-the-art VLMs on four benchmark image datasets captured in diverse environments. Our results offer insight into the internal reasoning of VLMs and highlight their strengths, limitations, and potential societal risks. Our findings indicate that current VLMs perform poorly on generic street-level images yet achieve notably high accuracy (61\%) on images resembling social media content, raising significant and urgent privacy concerns.

📄 PDF Abstract BibTeX arXiv:2508.19967

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mapping Public Perception of Artificial Intelligence: Expectations, Risk-Benefit Tradeoffs, and Value As Determinants for Societal Acceptance

2024-11-28 · Philipp Brauner, Felix Glawe, Gian Luca Liehner, Luisa Vervier 외

Understanding public perception of artificial intelligence (AI) and the tradeoffs between potential risks and benefits is crucial, as these perceptions might shape policy decisions, influence innovation trajectories for …

Autonomous Driving

Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models

2025-04-27 · Weidi Luo, Tianyu Lu, Qiming Zhang, Xiaogeng Liu 외

Recent advances in multi-modal large reasoning models (MLRMs) have shown significant ability to interpret complex visual content. While these models enable impressive reasoning capabilities, they also introduce novel and…

Visual ReasoningWorld Knowledge

Image-Based Geolocation Using Large Vision-Language Models

2024-08-18 · Yi Liu, Junchen Ding, Gelei Deng, Yuekang Li 외

Geolocation is now a vital aspect of modern life, offering numerous benefits but also presenting serious privacy concerns. The advent of large vision-language models (LVLMs) with advanced image-processing capabilities in…

Evaluating Precise Geolocation Inference Capabilities of Vision Language Models

2025-02-20 · Neel Jay, Hieu Minh Nguyen, Trung Dung Hoang, Jacob Haimes

The prevalence of Vision-Language Models (VLMs) raises important questions about privacy in an era where visual information is increasingly available. While foundation VLMs demonstrate broad knowledge and learned capabil…

AI Awareness

2025-04-25 · Xiaojian Li, Haoyuan Shi, Rongwu Xu, Wei Xu

Recent breakthroughs in artificial intelligence (AI) have brought about increasingly capable systems that demonstrate remarkable abilities in reasoning, language understanding, and problem-solving. These advancements hav…

Safety Alignment