paper-with-me

홈 › Papers

X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization

2026-08-17 · Zichao Zeng, Weijia Fan, Yufan Chen, June Moh Goo, Junwei Zheng, Ruiping Liu, Kunyu Peng, Jiaming Zhang, Rainer Stiefelhagen, Jan Boehm arxiv

Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. However, CVG approaches rely on fixed-length inputs and post-hoc refinement, hindering online-oriented localization under partial or dynamic observations. In this work, we formulate Progressive Cross-view Video Geo-localization (PCVG) as a deployment-oriented extension and evaluation protocol of CVG, enabling localization under varying temporal budgets, prefix-based inference, random-start evaluation, and long-range localization with interruptions. To explore PCVG, we introduce X$^2$Localizer, a cross-grained alignment framework that jointly supervises global prefix-to-aerial retrieval and token-aggregated frame--aerial-tile matching with a budget-dependent asymmetric objective. Furthermore, we introduce a Sliding-Window Re-Localization (SWRL) strategy that dynamically refreshes candidate regions for failure recovery and long-range deployment without full-sequence reprocessing. Extensive experiments show that X$^2$Localizer preserves conventional full-video performance, with marginal gains of +0.1 Recall@1 and +0.3 Recall@10, while substantially improving early localization. In the challenging single-frame setting, X$^2$Localizer improves coarse retrieval by +4.7 Recall@1 and +11.5 Recall@10 over the previous state-of-the-art method. With SWRL, our approach further enables robust progressive localization under random-start and long-distance scenarios, narrowing the gap between benchmark evaluation and real-world deployment.

📄 PDF Abstract BibTeX arXiv:2608.16658

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Progressive Adversarial Networks for Fine-Grained Domain Adaptation

2020-06-01 · CVPR 2020 6 · Sinan Wang, Xinyang Chen, Yunbo Wang, Mingsheng Long 외

Fine-grained visual categorization has long been considered as an important problem, however, its real application is still restricted, since precisely annotating a large fine-grained image dataset is a laborious task an…

Domain AdaptationFine-Grained Visual Categorization

UPDA: Unsupervised Progressive Domain Adaptation for No-Reference Point Cloud Quality Assessment

2026-02-12 · Bingxu Xie, Fang Zhou, Jincan Wu, Yonghui Liu 외 arxiv

While no-reference point cloud quality assessment (NR-PCQA) approaches have achieved significant progress over the past decade, their performance often degrades substantially when a distribution gap exists between the tr…

Point Cloud Quality AssessmentDomain Adaptation

CLAMP: Contrastive Learning with Adaptive Multi-loss and Progressive Fusion for Multimodal Aspect-Based Sentiment Analysis

2025-07-21 · Xiaoqiang He arxiv

Multimodal aspect-based sentiment analysis(MABSA) seeks to identify aspect terms within paired image-text data and determine their fine grained sentiment polarities, representing a fundamental task for improving the effe…

Contrastive LearningSentiment Analysis

Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models

2025-07-31 · Yassine Jamaa, Badr AlKhamissi, Satrajit Ghosh, Martin Schrimpf arxiv

This work adapts a neuroscientific contrast localizer to pinpoint causally relevant units for Theory of Mind (ToM) and mathematical reasoning tasks in large language models (LLMs) and vision-language models (VLMs). Acros…

Mathematical Reasoning

Localizing by Describing: Attribute-Guided Attention Localization for Fine-Grained Recognition

2016-05-20 · Xiao Liu, Jiang Wang, Shilei Wen, Errui Ding 외

A key challenge in fine-grained recognition is how to find and represent discriminative local regions. Recent attention models are capable of learning discriminative region localizers only from category labels with reinf…

Attributereinforcement-learningReinforcement LearningReinforcement Learning (RL)