paper-with-me

홈 › Papers

SceneGraphLoc: Cross-Modal Coarse Visual Localization on 3D Scene Graphs

2024-03-30 · Yang Miao, Francis Engelmann, Olga Vysotska, Federico Tombari, Marc Pollefeys, Dániel Béla Baráth

We introduce a novel problem, i.e., the localization of an input image within a multi-modal reference map represented by a database of 3D scene graphs. These graphs comprise multiple modalities, including object-level point clouds, images, attributes, and relationships between objects, offering a lightweight and efficient alternative to conventional methods that rely on extensive image databases. Given the available modalities, the proposed method SceneGraphLoc learns a fixed-sized embedding for each node (i.e., representing an object instance) in the scene graph, enabling effective matching with the objects visible in the input query image. This strategy significantly outperforms other cross-modal methods, even without incorporating images into the map embeddings. When images are leveraged, SceneGraphLoc achieves performance close to that of state-of-the-art techniques depending on large image databases, while requiring three orders-of-magnitude less storage and operating orders-of-magnitude faster. The code will be made public.

📄 PDF Abstract BibTeX arXiv:2404.00469

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Localization

Similar Papers 제목 키워드 기반

Cross-Modal Visual Relocalization in Prior LiDAR Maps Utilizing Intensity Textures

2024-12-02 · Qiyuan Shen, Hengwang Zhao, Weihao Yan, Chunxiang Wang 외

Cross-modal localization has drawn increasing attention in recent years, while the visual relocalization in prior LiDAR maps is less studied. Related methods usually suffer from inconsistency between the 2D texture and 3…

3D geometryPose EstimationRetrieval

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

2024-12-17 · Ziheng Zhou, Jinxing Zhou, Wei Qian, Shengeng Tang 외

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene unde…

audio-visual event localizationaudio-visual learningScene Understanding

Multiple Sound Sources Localization from Coarse to Fine

2020-07-13 · ECCV 2020 8 · Rui Qian, Di Hu, Heinrich Dinkel, Mengyue Wu 외

How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this problem, we develop a two-stage audiovisual le…

From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching

2026-07-29 · Shiyong Meng, Bolei Chen, Ping Zhong, Yang Wan 외 arxiv

Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching egocentric images against minimalist structural maps. However, due to cross-modal information asymmetry and rep…

Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake Localization

2026-05-22 · Jiayu Xiong, Jing Wang, Qi Zhang, Wanlong Wang 외 arxiv

Audio-visual deepfake localization demands interval-level outputs that serve as temporal evidence. Despite recent progress, symmetric fusion under single-sided or asynchronous forgeries propagates cross-modal noise, degr…