paper-with-me

Papers

Multi-View Deep Learning for Consistent Semantic Mapping with RGB-D Cameras

2017-03-26 · Lingni Ma, Jörg Stückler, Christian Kerl, Daniel Cremers

Visual scene understanding is an important capability that enables robots to purposefully act in their environment. In this paper, we propose a novel approach to object-class segmentation from multiple RGB-D views using deep learning. We train a deep neural network to predict object-class semantics that is consistent from several view points in a semi-supervised way. At test time, the semantics predictions of our network can be fused more consistently in semantic keyframe maps than predictions of a network trained on individual views. We base our network architecture on a recent single-view deep learning approach to RGB and depth fusion for semantic object-class segmentation and enhance it with multi-scale loss minimization. We obtain the camera trajectory using RGB-D SLAM and warp the predictions of RGB-D images into ground-truth annotated frames in order to enforce multi-view consistency during training. At test time, predictions from multiple views are fused into keyframes. We propose and analyze several methods for enforcing multi-view consistency during training and testing. We evaluate the benefit of multi-view consistency training and demonstrate that pooling of deep features and fusion over multiple views outperforms single-view baselines on the NYUDv2 benchmark for semantic segmentation. Our end-to-end trained network achieves state-of-the-art performance on the NYUDv2 dataset in single-view segmentation as well as multi-view semantic fusion.

📄 PDF Abstract BibTeX arXiv:1703.08866

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

Cross-view Transformers for real-time Map-view Semantic Segmentation

2022-05-05 · CVPR 2022 1 · Brady Zhou, Philipp Krähenbühl

We present cross-view transformers, an efficient attention-based model for map-view semantic segmentation from multiple cameras. Our architecture implicitly learns a mapping from individual camera views into a canonical …

Bird's-Eye View Semantic SegmentationSegmentationSemantic Segmentation

OneBEV: Using One Panoramic Image for Bird's-Eye-View Semantic Mapping

2024-09-20 · Jiale Wei, Junwei Zheng, Ruiping Liu, Jie Hu 외

In the field of autonomous driving, Bird's-Eye-View (BEV) perception has attracted increasing attention in the community since it provides more comprehensive information compared with pinhole front-view images and panora…

Autonomous DrivingMamba

Loop Closure Detection Based on Object-level Spatial Layout and Semantic Consistency

2023-04-11 · Xingwu Ji, Peilin Liu, Haochen Niu, Xiang Chen 외

Visual simultaneous localization and mapping (SLAM) systems face challenges in detecting loop closure under the circumstance of large viewpoint changes. In this paper, we present an object-based loop closure detection me…

Graph MatchingLoop Closure DetectionObjectSimultaneous Localization and Mapping

On the Advantages of Multiple Stereo Vision Camera Designs for Autonomous Drone Navigation

2021-05-26 · Rui Pimentel de Figueiredo, Jakob Grimm Hansen, Jonas Le Fevre, Martim Brandão 외

In this work we showcase the design and assessment of the performance of a multi-camera UAV, when coupled with state-of-the-art planning and mapping algorithms for autonomous navigation. The system leverages state-of-the…

Autonomous NavigationDrone navigation

TouchMap-OR: Multi-View 3D Mapping of Hand-Surface Contacts

2026-05-17 · Sophokles Ktistakis, Rui Wang, Bastian Grande, Hugo Sax arxiv

Hand-surface interactions between clinicians, patients, and medical equipment play a central role in pathogen transmission during medical procedures. However, these interactions remain largely unobserved, as current infe…