paper-with-me

홈 › Papers

Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation

2024-03-28 · Zhongliang Zhou, Jielu Zhang, Zihan Guan, Mengxuan Hu, Ni Lao, Lan Mu, Sheng Li, Gengchen Mai

Geolocating precise locations from images presents a challenging problem in computer vision and information retrieval.Traditional methods typically employ either classification, which dividing the Earth surface into grid cells and classifying images accordingly, or retrieval, which identifying locations by matching images with a database of image-location pairs. However, classification-based approaches are limited by the cell size and cannot yield precise predictions, while retrieval-based systems usually suffer from poor search quality and inadequate coverage of the global landscape at varied scale and aggregation levels. To overcome these drawbacks, we present Img2Loc, a novel system that redefines image geolocalization as a text generation task. This is achieved using cutting-edge large multi-modality models like GPT4V or LLaVA with retrieval augmented generation. Img2Loc first employs CLIP-based representations to generate an image-based coordinate query database. It then uniquely combines query results with images itself, forming elaborate prompts customized for LMMs. When tested on benchmark datasets such as Im2GPS3k and YFCC4k, Img2Loc not only surpasses the performance of previous state-of-the-art models but does so without any model training.

📄 PDF Abstract BibTeX arXiv:2403.19584

Code (1)

Douglas2Code/Img2Loc 공식 구현 pytorch

Tasks

RetrievalRetrieval-augmented GenerationText Generation

Similar Papers 제목 키워드 기반

Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization

2026-04-17 · Siddhant Bharadwaj, Ashish Vashist, Fahimul Aleem, Shruti Vyas arxiv

Image geolocalization has traditionally been addressed through retrieval-based place recognition or geometry-based visual localization pipelines. Recent advances in Vision-Language Models (VLMs) have demonstrated strong …

Multimodal ReasoningVisual LocalizationImage Matching

Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization

2023-02-01 · Lukas Haas, Silas Alberti, Michal Skreta

Image geolocalization is the challenging task of predicting the geographic coordinates of origin for a given photo. It is an unsolved problem relying on the ability to combine visual clues with general knowledge about th…

Generalized Zero-Shot LearningMeta-LearningPhoto geolocation estimationZero-Shot Learning

Revisiting IM2GPS in the Deep Learning Era

2017-05-13 · ICCV 2017 10 · Nam Vo, Nathan Jacobs, James Hays

Image geolocalization, inferring the geographic location of an image, is a challenging computer vision problem with many potential applications. The recent state-of-the-art approach to this problem is a deep image classi…

ClassificationDeep LearningDensity EstimationGeneral Classification+5

CityGuessr: City-Level Video Geo-Localization on a Global Scale

2024-11-10 · Parth Parag Kulkarni, Gaurav Kumar Nayak, Mubarak Shah

Video geolocalization is a crucial problem in current times. Given just a video, ascertaining where it was captured from can have a plethora of advantages. The problem of worldwide geolocalization has been tackled before…

geo-localization

G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models

2024-05-23 · Pengyue Jia, Yiding Liu, Xiaopeng Li, Yuhao Wang 외

Worldwide geolocalization aims to locate the precise location at the coordinate level of photos taken anywhere on the Earth. It is very challenging due to 1) the difficulty of capturing subtle location-aware visual seman…

Photo geolocation estimationRAGRetrievalRetrieval-augmented Generation