paper-with-me

Papers

GaGA: Towards Interactive Global Geolocation Assistant

2024-12-12 · Zhiyang Dou, Zipeng Wang, Xumeng Han, Guorong Li, Zhipei Huang, Zhenjun Han

Global geolocation, which seeks to predict the geographical location of images captured anywhere in the world, is one of the most challenging tasks in the field of computer vision. In this paper, we introduce an innovative interactive global geolocation assistant named GaGA, built upon the flourishing large vision-language models (LVLMs). GaGA uncovers geographical clues within images and combines them with the extensive world knowledge embedded in LVLMs to determine the geolocations while also providing justifications and explanations for the prediction results. We further designed a novel interactive geolocation method that surpasses traditional static inference approaches. It allows users to intervene, correct, or provide clues for the predictions, making the model more flexible and practical. The development of GaGA relies on the newly proposed Multi-modal Global Geolocation (MG-Geo) dataset, a comprehensive collection of 5 million high-quality image-text pairs. GaGA achieves state-of-the-art performance on the GWS15k dataset, improving accuracy by 4.57% at the country level and 2.92% at the city level, setting a new benchmark. These advancements represent a significant leap forward in developing highly accurate, interactive geolocation systems with global applicability.

📄 PDF Abstract BibTeX arXiv:2412.08907

Code (0)

등록된 구현이 없습니다.

Tasks

World Knowledge

Similar Papers 제목 키워드 기반

Real-time Classification, Geolocation and Interactive Visualization of COVID-19 Information Shared on Social Media to Better Understand Global Developments

2020-12-01 · EMNLP (NLP-COVID19) 2020 12 · Andrei Mircea

As people communicate on social media during COVID-19, it can be an invaluable source of useful and up-to-date information. However, the large volume and noise-to-signal ratio of social media can make this impractical. W…

Classification

Learning to Wander: Improving the Global Image Geolocation Ability of LMMs via Actionable Reasoning

2026-03-11 · Yushuo Zheng, Huiyu Duan, Zicheng Zhang, Xiaohong Liu 외 arxiv

Geolocation, the task of identifying the geographic location of an image, requires abundant world knowledge and complex reasoning abilities. Though advanced large multimodal models (LMMs) have shown superior aforemention…

G^3: Geolocation via Guidebook Grounding

2022-11-28 · Grace Luo, Giscard Biamby, Trevor Darrell, Daniel Fried 외

We demonstrate how language can improve geolocation: the task of predicting the location where an image was taken. Here we study explicit knowledge from human-written guidebooks that describe the salient and class-discri…

Comparing Traditional and LLM-based Search for Image Geolocation

2024-01-18 · Albatool Wazzan, Stephen MacNeil, Richard Souvenir

Web search engines have long served as indispensable tools for information retrieval; user behavior and query formulation strategies have been well studied. The introduction of search engines powered by large language mo…

Conversational SearchInformation RetrievalNatural Language QueriesRetrieval

VideoGigaGAN: Towards Detail-rich Video Super-Resolution

2024-04-18 · CVPR 2025 1 · Yiran Xu, Taesung Park, Richard Zhang, Yang Zhou 외

Video super-resolution (VSR) approaches have shown impressive temporal consistency in upsampled videos. However, these approaches tend to generate blurrier results than their image counterparts as they are limited in the…

Super-ResolutionVideo Super-Resolution