Indoor Semantic Scene Understanding using Multi-modality Fusion
Seamless Human-Robot Interaction is the ultimate goal of developing service robotic systems. For this, the robotic agents have to understand their surroundings to better complete a given task. Semantic scene understanding allows a robotic agent to extract semantic knowledge about the objects in the environment. In this work, we present a semantic scene understanding pipeline that fuses 2D and 3D detection branches to generate a semantic map of the environment. The 2D mask proposals from state-of-the-art 2D detectors are inverse-projected to the 3D space and combined with 3D detections from point segmentation networks. Unlike previous works that were evaluated on collected datasets, we test our pipeline on an active photo-realistic robotic environment - BenchBot. Our novelty includes rectification of 3D proposals using projected 2D detections and modality fusion based on object size. This work is done as part of the Robotic Vision Scene Understanding Challenge (RVSU). The performance evaluation demonstrates that our pipeline has improved on baseline methods without significant computational bottleneck.
Code (0)
등록된 구현이 없습니다.
Tasks
Scene UnderstandingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Towards Multimodal Multitask Scene Understanding Models for Indoor Mobile Agents
The perception system in personalized mobile agents requires developing indoor scene understanding models, which can understand 3D geometries, capture objectiveness, analyze human behaviors, etc. Nonetheless, this direct…
3D Object DetectionAutonomous DrivingComputational EfficiencyDepth Completion+8Shallow2Deep: Indoor Scene Modeling by Single Image Understanding
Dense indoor scene modeling from 2D images has been bottlenecked due to the absence of depth information and cluttered occlusions. We present an automatic indoor scene modeling approach using deep features from neural ne…
3D geometryglobal-optimizationRelation NetworkScene UnderstandingUniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
3D open-vocabulary scene understanding aims to recognize arbitrary novel categories beyond the base label space. However, existing works not only fail to fully utilize all the available modal information in the 3D domain…
Instance SegmentationScene UnderstandingSemantic SegmentationRoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing
Generating controllable indoor scenes is fundamental to applications in game development, architectural visualization, and embodied AI. However, existing approaches either support a limited input modalities or rely on im…
Indoor Scene SynthesisScene GenerationSemantic ParsingSmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these mod…
Multimodal ReasoningScene Understanding