paper-with-me

홈 › Papers

VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimation

2026-03-13 · Juhye Park, Wooju Lee, Dasol Hong, Changki Sung, Youngwoo Seo, Dongwan Kang, Hyun Myung arxiv

Accurate global localization is critical for autonomous driving and robotics, but GNSS-based approaches often degrade due to occlusion and multipath effects. As an emerging alternative, cross-view pose estimation predicts the 3-DoF camera pose corresponding to a ground-view image with respect to a geo-referenced satellite image. However, existing methods struggle to bridge the significant viewpoint gap between the ground and satellite views mainly due to limited spatial correspondences. We propose a novel cross-view pose estimation method that constructs view-invariant representations through dual-axis transformation (VIRD). VIRD first applies a polar transformation to the satellite view to facilitate horizontal correspondence, then uses context-enhanced positional attention on the ground and polar-transformed satellite features to mitigate vertical misalignment, explicitly bridging the viewpoint gap. To further strengthen view invariance, we introduce a view-reconstruction loss that encourages the derived representations to reconstruct the original and cross-view images. Experiments on the KITTI and VIGOR datasets demonstrate that VIRD outperforms the state-of-the-art methods without orientation priors, reducing median position and orientation errors by 50.7% and 76.5% on KITTI, and 18.0% and 46.8% on VIGOR, respectively.

📄 PDF Abstract BibTeX arXiv:2603.12918

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingPose Estimation

Similar Papers 제목 키워드 기반

VIRDOCD: a VIRtual DOCtor to Predict Dengue Fatality

2021-04-29 · Amit K Chattopadhyay, Subhagata Chattopadhyay

Clinicians make routine diagnosis by scrutinizing patients' medical signs and symptoms, a skill popularly referred to as "Clinical Eye". This skill evolves through trial-and-error and improves with time. The success of t…

BIG-bench Machine LearningDecision Makingregression

ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation

2026-09-24 · Zichong Meng, Chongjian Ge, Chun-Hao P. Huang, Yang Zhou 외 hf

Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both a large pretrained …

Image GenerationVideo Generation

VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming

2025-10-02 · Duy Nguyen, Dat Nguyen arxiv

Existing UDA pipelines fine-tune already well-trained backbone parameters for every new source-and-target pair, resulting in the number of training parameters and storage memory growing linearly with each new pair, and a…

Unsupervised Domain Adaptation

LEVIRDet: A Million-Scale 159-Category Dataset and Foundation Model for Universal Remote Sensing Object Detection

2026-06-24 · Qinzhe Yang, Dongyu Wang, Haohan Niu, Jia Xu 외 arxiv

Remote sensing object detection has advanced rapidly with the development of large-scale benchmarks and modern detection architectures. However, existing datasets and detectors remain fragmented. Most benchmarks focus on…

Domain GeneralizationObject Detection

Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learning

2025-01-01 · CVPR 2025 1 · Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman

Egocentric and exocentric perspectives of human action differ significantly, yet overcoming this extreme viewpoint gap is critical for applications in augmented reality and robotics. We propose ViewpointRosetta, an a…

Action RecognitionContrastive LearningRepresentation Learning