paper-with-me

홈 › Papers

When Vision Misleads, Let Location Speak: A Worldwide Image Geo-Localization Method via Location Attention Mechanism and Large Multimodal Models

2026-06-08 · Junchao Cui, Wenqi Shi, Xuanzi Ma, Nan Wu, Shaoyong Du, Xiangyang Luo arxiv

Worldwide image geo-localization aims to determine the capture location of an image on a global scale. Existing methods often mislocalize images by matching them to visually similar scenes from different geographic regions, which limits reliability in practical applications. To address this issue, we propose TransGeoCLIP, a novel retrieval-based framework that integrates a location attention mechanism and large multimodal models (LMMs). Using the Transformer encoder with location attention to encode GPS coordinates, TransGeoCLIP can effectively distinguish geographic features among visually similar images. The framework consists of two stages: 1) Retrieval database construction, which employs Transformers equipped with location attention mechanisms to encode labeled GPS coordinates and enhance location semantics, subsequently enables joint image-text-GPS embedding through CLIP; 2) Retrieval-augmented inference, which leverages LMMs to infer the final image location prediction from retrieved database results. Extensive experimental results on diverse datasets, including IM2GPS, IM2GPS3k, YFCC4k, and YFCC26k, demonstrate that TransGeoCLIP significantly enhances localization performance for visually similar images. Particularly, street-level localization accuracy (within 1 km error) is substantially improved, surpassing state-of-the-art methods by 1.5%, 1.07%, 7.18%, and 9.75% on these benchmarks, respectively.

📄 PDF Abstract BibTeX arXiv:2606.08918

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Conditional Teacher-Student Learning

2019-04-28 · Zhong Meng, Jinyu Li, Yong Zhao, Yifan Gong

The teacher-student (T/S) learning has been shown to be effective for a variety of problems such as domain adaptation and model compression. One shortcoming of the T/S learning is that a teacher model, not always perfect…

Domain AdaptationModel Compression

G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models

2024-05-23 · Pengyue Jia, Yiding Liu, Xiaopeng Li, Yuhao Wang 외

Worldwide geolocalization aims to locate the precise location at the coordinate level of photos taken anywhere on the Earth. It is very challenging due to 1) the difficulty of capturing subtle location-aware visual seman…

Photo geolocation estimationRAGRetrievalRetrieval-augmented Generation

When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models

2026-05-27 · Jungwon Park, Jimyeong Kim, Jungmin Ko, Nojun Kwak 외 arxiv

Diffusion language models decode text by iteratively denoising masked token sequences, making the choice of which positions to decode a central inference-time decision. Most training-free decoding strategies use model co…

Joint speaker diarisation and tracking in switching state-space model

2021-09-23 · Jeremy H. M. Wong, Yifan Gong

Speakers may move around while diarisation is being performed. When a microphone array is used, the instantaneous locations of where the sounds originated from can be estimated, and previous investigations have shown tha…

Cross modal video representations for weakly supervised active speaker localization

2020-03-09 · Rahul Sharma, Krishna Somandepalli, Shrikanth Narayanan

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, …

Action DetectionActive Speaker LocalizationActivity DetectionEvent Detection