paper-with-me

홈 › Papers

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale

2026-04-14 · Parth Parag Kulkarni, Rohit Gupta, Prakash Chandra Chhipa, Mubarak Shah arxiv

The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensics, social media, and exploration. Existing classification-based approaches operate at a coarse city-level granularity and fail to capture fine-grained details, while image retrieval methods are impractical on a global scale due to the need for extensive image galleries which are infeasible to compile. Comparatively, constructing a gallery of GPS coordinates is straightforward and inexpensive. We propose VidTAG, a dual-encoder framework that performs frame-to-GPS retrieval using both self-supervised and language-aligned features. To address temporal inconsistencies in video predictions, we introduce the TempGeo module, which aligns frame embeddings, and the GeoRefiner module, an encoder-decoder architecture that refines GPS features using the aligned frame embeddings. Evaluations on Mapillary (MSLS) and GAMa datasets demonstrate our model's ability to generate temporally consistent trajectories and outperform baselines, achieving a 20% improvement at the 1 km threshold over GeoCLIP. We also beat current State-of-the-Art by 25% on global coarse grained video geolocalization (CityGuessr68k). Our approach enables fine-grained video geolocalization and lays a strong foundation for future research. More details on the project webpage: https://parthpk.github.io/vidtag_webpage/

📄 PDF Abstract BibTeX arXiv:2604.12159

Code (0)

등록된 구현이 없습니다.

Tasks

Image Retrieval

Similar Papers 제목 키워드 기반

CityGuessr: City-Level Video Geo-Localization on a Global Scale

2024-11-10 · Parth Parag Kulkarni, Gaurav Kumar Nayak, Mubarak Shah

Video geolocalization is a crucial problem in current times. Given just a video, ascertaining where it was captured from can have a plethora of advantages. The problem of worldwide geolocalization has been tackled before…

geo-localization

Map-Based Temporally Consistent Geolocalization through Learning Motion Trajectories

2020-10-13 · Bing Zha, Alper Yilmaz

In this paper, we propose a novel trajectory learning method that exploits motion trajectories on topological map using recurrent neural network for temporally consistent geolocalization of object. Inspired by human's ab…

Visual Odometry

Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation

2026-05-24 · Shuyuan Tu, Qi Tian, Zihan Yang, Yue Wu 외 arxiv

Current open-source diffusion models struggle to generate stable and synchronized audio-visual content, particularly in scenarios demanding complex semantic reasoning. The root cause is that existing methods rely on coar…

Audio GenerationVideo Denoising

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

2024-12-19 · Moayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov 외

We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioni…

Video GenerationVideo Synchronization

GAMa: Cross-view Video Geo-localization

2022-07-06 · Shruti Vyas, Chen Chen, Mubarak Shah

The existing work in cross-view geo-localization is based on images where a ground panorama is matched to an aerial image. In this work, we focus on ground videos instead of images which provides additional contextual cu…

geo-localization