paper-with-me

홈 › Papers

SurroundNEXO: Ego-Centric Metric Bridging for Spatially Consistent Geometry in Autonomous Driving

2026-06-15 · Shuai Yuan, Runxi Tang, Yuzhou Ji, Fudong Ge, Hanshi Wang, Yifei Wang, Xianming Zeng, Jianyun Xu, Xingliang Liu, Yanfeng Wang, Zhipeng Zhang arxiv

Modern autonomous driving depends on accurate metric 3D understanding for perception, reconstruction, and planning, which in turn requires reliable multi-camera depth prediction. However, the outward-facing nature of vehicle-mounted surround-view camera rigs inherently limits visual overlap across views, challenging the correspondence-based assumptions that underpin conventional multi-view geometry. To bridge this gap, we present SurroundNEXO, named after the Spanish word nexo for a geometric link, a low-overlap multi-camera metric depth framework that grounds cross-view reasoning in ego-centric geometry rather than dense visual correspondences. Instead of directly enforcing early global fusion, SurroundNEXO first assigns image tokens globally comparable ego-frame viewing directions through Ego-Ray Positional Encoding, then uses sparse LiDAR measurements as metric anchors to propagate absolute scale cues, and finally expands feature interaction progressively from view-local modeling to decomposed spatio-temporal reasoning and global integration. This design enables metric-scale depth prediction with improved spatial consistency across weakly overlapping cameras. Across low-overlap autonomous driving benchmarks, including NuScenes, Waymo and DDAD, SurroundNEXO reduces single-view error by 33.2%, improves cross-view consistency by 10.5%, and enhances metric reconstruction quality by 25.6% compared with SOTA methods. It further remains robust under extremely sparse depth prompts and exhibits strong zero-shot generalization to unseen camera layouts.

📄 PDF Abstract BibTeX arXiv:2606.16960

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationAutonomous Driving

Similar Papers 제목 키워드 기반

EgoX: Egocentric Video Generation from a Single Exocentric Video

2025-12-09 · Taewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park 외 arxiv

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibili…

Video Generation

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs

2026-05-10 · Bo Gu, Zhikang Zhang, Zizhuang Wei, Zhenyuan Chen 외 arxiv

Recent multimodal large language models (MLLMs) have made remarkable progress in visual understanding and language-based reasoning, yet they lack a persistent world-centered representation for spatially consistent reason…

Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation

2026-02-05 · Hengyi Wang, Ruiqiang Zhang, Chang Liu, Guanjie Wang 외 arxiv

With the rising need for spatially grounded tasks such as Vision-Language Navigation/Action, allocentric perception capabilities in Vision-Language Models (VLMs) are receiving growing focus. However, VLMs remain brittle …

Vision-Language NavigationSpatial Reasoning

From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning

2026-02-03 · Hyun Seok Seong, WonJun Moon, Jae-Pil Heo arxiv

Unsupervised object-centric learning models, particularly slot-based architectures, have shown great promise in decomposing complex scenes. However, their reliance on reconstruction-based training creates a fundamental c…

Representation Learning

COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

2025-06-16 · Yining Shi, Kun Jiang, Qiang Meng, Ke Wang 외

World models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data. Existing methods struggle to disentangle ego-vehicle motion (perspective shifts) from scene evolvement (age…

Autonomous DrivingRepresentation Learning