paper-with-me

홈 › Papers

Scene Grounding In the Wild

2026-03-27 · Tamir Cohen, Leo Segre, Shay Shomer-Chai, Shai Avidan, Hadar Averbuch-Elor arxiv

Reconstructing accurate 3D models of large-scale real-world scenes from unstructured, in-the-wild imagery remains a core challenge in computer vision, especially when the input views have little or no overlap. In such cases, existing reconstruction pipelines often produce multiple disconnected partial reconstructions or erroneously merge non-overlapping regions into overlapping geometry. In this work, we propose a framework that grounds each partial reconstruction to a complete reference model of the scene, enabling globally consistent alignment even in the absence of visual overlap. We obtain reference models from dense, geospatially accurate pseudo-synthetic renderings derived from Google Earth Studio. These renderings provide full scene coverage but differ substantially in appearance from real-world photographs. Our key insight is that, despite this significant domain gap, both domains share the same underlying scene semantics. We represent the reference model using 3D Gaussian Splatting, augmenting each Gaussian with semantic features, and formulate alignment as an inverse feature-based optimization scheme that estimates a global 6DoF pose and scale while keeping the reference model fixed. Furthermore, we introduce the WikiEarth dataset, which registers existing partial 3D reconstructions with pseudo-synthetic reference models. We demonstrate that our approach consistently improves global alignment when initialized with various classical and learning-based pipelines, while mitigating failure modes of state-of-the-art end-to-end models.

📄 PDF Abstract BibTeX arXiv:2603.26584

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural Language

2023-04-12 · Zhenxiang Lin, Xidong Peng, Peishan Cong, Ge Zheng 외

We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images and 3D LiDAR point clouds. We present a…

3D visual groundingAutonomous DrivingObject LocalizationVisual Grounding

AffordanceLLM: Grounding Affordance from Vision Language Models

2024-01-12 · Shengyi Qian, Weifeng Chen, Min Bai, Xiong Zhou 외

Affordance grounding refers to the task of finding the area of an object with which one can interact. It is a fundamental but challenging task, as a successful solution requires the comprehensive understanding of a scene…

Human-Object Interaction DetectionObject

Optimizing Battery and Line Undergrounding Investments for Transmission Systems under Wildfire Risk Scenarios: A Benders Decomposition Approach

2025-06-07 · Ryan Piansky, Rahul K. Gupta, Daniel K. Molzahn

With electric power infrastructure posing an increasing risk of igniting wildfires under continuing climate change, utilities are frequently de-energizing power lines to mitigate wildfire ignition risk, which can cause l…

EgoSim: Egocentric World Simulator for Embodied Interaction Generation

2026-04-01 · Jinkun Hao, Mingda Jia, Ruiyan Wang, Hongrui Zhu 외 arxiv

We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation. Existing egocentric s…

Point Clouds

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents

2026-05-19 · Bingnan Liu, Chenhang Cui, Rui Huang, Jiani Luo 외 arxiv

We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professio…

Visual Grounding