paper-with-me

Papers

More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

2026-07-14 · Mao Chen, Xiangkai Zhang, Zhiyong Liu, Chuankai Liu, Xu Yang arxiv

Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a promising pathway toward this ability, as it requires a model to align ground-view imagery with geo-referenced satellite-view imagery despite drastic appearance changes to estimate camera poses. Recent visual foundation models have made this long-standing localization problem increasingly feasible by providing rich 2D representations for cross-view matching. However, we argue that cross-view localization should not be viewed merely as 2D matching or pose estimation. In this work, we revisit cross-view localization as more than pose estimation and investigate how it can help the model develop consistent cross-view understanding under extreme viewpoint changes, including stable semantics, reliable structure, and transferable geometry. We identify three key limitations of existing methods that prevent them from achieving this. They usually lack explicit 3D grounding, rely on strict point-wise matching that can weaken semantic consistency, and learn from an absolute objective that provides limited guidance for geometric reasoning. To address these limitations, we propose CROSS, a unified cross-view localization framework built upon 3D-grounded alignment, structure-aware matching, and hypothesis ranking. This formulation makes structure learning an intrinsic requirement, encourages semantic representations to remain stable, and enables the model to acquire transferable geometry. Extensive experiments on the KITTI and VIGOR datasets show that CROSS achieves state-of-the-art performance in cross-view localization. More importantly, CROSS effectively learns stable semantics, reliable structure, and transferable geometry across extremely different viewpoints.

📄 PDF Abstract BibTeX arXiv:2607.12429

Code (0)

등록된 구현이 없습니다.

Tasks

Pose Estimation

Similar Papers 제목 키워드 기반

VoxFormer: Sparse Voxel Transformer for Camera-based 3D Semantic Scene Completion

2023-02-23 · CVPR 2023 1 · Yiming Li, Zhiding Yu, Christopher Choy, Chaowei Xiao 외

Humans can easily imagine the complete 3D geometry of occluded objects and scenes. This appealing ability is vital for recognition and understanding. To enable such capability in AI systems, we propose VoxFormer, a Trans…

3D geometry3D Semantic Scene Completion3D Semantic Scene Completion from a single RGB imageDepth Estimation+1

Shape from Semantics: 3D Shape Generation from Multi-View Semantics

2025-02-01 · Liangchen Li, Caoliwen Wang, Yuqi Zhou, Bailin Deng 외

We propose ``Shape from Semantics'', which is able to create 3D models whose geometry and appearance match given semantics when observed from different views. Traditional ``Shape from X'' tasks usually use visual input (…

3D geometry3D Shape GenerationImage RestorationVideo Generation

CUS-GS: A Compact Unified Structured Gaussian Splatting Framework for Multimodal Scene Representation

2025-11-22 · Yuhang Ming, Chenxin Fang, Xingyuan Yu, Fan Zhang 외 arxiv

Recent advances in Gaussian Splatting based 3D scene representation have shown two major trends: semantics-oriented approaches that focus on high-level understanding but lack explicit 3D geometry modeling, and structure-…

VAST: The Valence-Assessing Semantics Test for Contextualizing Language Models

2022-03-14 · Robert Wolfe, Aylin Caliskan

VAST, the Valence-Assessing Semantics Test, is a novel intrinsic evaluation task for contextualized word embeddings (CWEs). VAST uses valence, the association of a word with pleasantness, to measure the correspondence of…

Word EmbeddingsWord Similarity

HyNeuralMap: Hyperbolic Mapping of Visual Semantics to Neural Hierarchies

2026-05-10 · Zihan Ma, Tian Xia, Kexin Wang, Xiao Li 외 arxiv

Understanding the intricate mappings between visual stimuli and neural responses is a fundamental challenge in cognitive neuroscience. While current approaches predominantly align images and functional magnetic resonance…

Representation LearningCross-Modal Retrieval