paper-with-me

Papers

Efficient Multi-Task RGB-D Scene Analysis for Indoor Environments

2022-07-10 · Daniel Seichter, Söhnke Benedikt Fischedick, Mona Köhler, Horst-Michael Groß

Semantic scene understanding is essential for mobile agents acting in various environments. Although semantic segmentation already provides a lot of information, details about individual objects as well as the general scene are missing but required for many real-world applications. However, solving multiple tasks separately is expensive and cannot be accomplished in real time given limited computing and battery capabilities on a mobile platform. In this paper, we propose an efficient multi-task approach for RGB-D scene analysis~(EMSANet) that simultaneously performs semantic and instance segmentation~(panoptic segmentation), instance orientation estimation, and scene classification. We show that all tasks can be accomplished using a single neural network in real time on a mobile platform without diminishing performance - by contrast, the individual tasks are able to benefit from each other. In order to evaluate our multi-task approach, we extend the annotations of the common RGB-D indoor datasets NYUv2 and SUNRGB-D for instance segmentation and orientation estimation. To the best of our knowledge, we are the first to provide results in such a comprehensive multi-task setting for indoor scene analysis on NYUv2 and SUNRGB-D.

📄 PDF Abstract BibTeX arXiv:2207.04526

Code (2)

tui-nicr/emsanet 공식 구현 pytorch
tui-nicr/nicr-scene-analysis-datasets pytorch

Tasks

Instance SegmentationPanoptic SegmentationScene ClassificationScene Classification (unified classes)Scene UnderstandingSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

IndoorR2X: Indoor Robot-to-Everything Coordination with LLM-Driven Planning

2026-03-20 · Fan Yang, Soumya Teotia, Shaunak A. Mehta, Prajit KrisshnaKumar 외 arxiv

Although robot-to-robot (R2R) communication improves indoor scene understanding beyond what a single robot can achieve, R2R alone cannot overcome partial observability without substantial exploration overhead or scaling …

Robot Task PlanningScene Understanding

ROOT: VLM based System for Indoor Scene Understanding and Beyond

2024-11-24 · Yonghui Wang, Shi-Yong Chen, Zhenxing Zhou, Siyi Li 외

Recently, Vision Language Models (VLMs) have experienced significant advancements, yet these models still face challenges in spatial hierarchical reasoning within indoor scenes. In this study, we introduce ROOT, a VLM-ba…

Scene GenerationScene Understanding

Towards Multimodal Multitask Scene Understanding Models for Indoor Mobile Agents

2022-09-27 · Yao-Hung Hubert Tsai, Hanlin Goh, Ali Farhadi, Jian Zhang

The perception system in personalized mobile agents requires developing indoor scene understanding models, which can understand 3D geometries, capture objectiveness, analyze human behaviors, etc. Nonetheless, this direct…

3D Object DetectionAutonomous DrivingComputational EfficiencyDepth Completion+8

IndoorCrowd: A Multi-Scene Dataset for Human Detection, Segmentation, and Tracking with an Automated Annotation Pipeline

2026-04-02 · Sebastian-Ion Nae, Radu Moldoveanu, Alexandra Stefania Ghita, Adina Magda Florea arxiv

Understanding human behaviour in crowded indoor environments is central to surveillance, smart buildings, and human-robot interaction, yet existing datasets rarely capture real-world indoor complexity at scale. We introd…

Multi-Object TrackingInstance Segmentation

Beyond Controlled Environments: 3D Camera Re-Localization in Changing Indoor Scenes

2020-08-05 · ECCV 2020 8 · Johanna Wald, Torsten Sattler, Stuart Golodetz, Tommaso Cavallari 외

Long-term camera re-localization is an important task with numerous computer vision and robotics applications. Whilst various outdoor benchmarks exist that target lighting, weather and seasonal changes, far less attentio…

Camera Relocalization