Towards Multimodal Multitask Scene Understanding Models for Indoor Mobile Agents
The perception system in personalized mobile agents requires developing indoor scene understanding models, which can understand 3D geometries, capture objectiveness, analyze human behaviors, etc. Nonetheless, this direction has not been well-explored in comparison with models for outdoor environments (e.g., the autonomous driving system that includes pedestrian prediction, car detection, traffic sign recognition, etc.). In this paper, we first discuss the main challenge: insufficient, or even no, labeled data for real-world indoor environments, and other challenges such as fusion between heterogeneous sources of information (e.g., RGB images and Lidar point clouds), modeling relationships between a diverse set of outputs (e.g., 3D object locations, depth estimation, and human poses), and computational efficiency. Then, we describe MMISM (Multi-modality input Multi-task output Indoor Scene understanding Model) to tackle the above challenges. MMISM considers RGB images as well as sparse Lidar points as inputs and 3D object detection, depth completion, human pose estimation, and semantic segmentation as output tasks. We show that MMISM performs on par or even better than single-task models; e.g., we improve the baseline 3D object detection results by 11.7% on the benchmark ARKitScenes dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Object DetectionAutonomous DrivingComputational EfficiencyDepth CompletionDepth EstimationObjectobject-detectionObject DetectionPose EstimationScene UnderstandingSemantic SegmentationTraffic Sign RecognitionSimilar Papers 제목 키워드 기반
Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning
Modern robot navigation systems encounter difficulties in diverse and complex indoor environments. Traditional approaches rely on multiple modules with small models or rule-based systems and thus lack adaptability to new…
Robot NavigationSelf-Supervised LearningVisual Place RecognitionMTANet: Multitask-Aware Network With Hierarchical Multimodal Fusion for RGB-T Urban Scene Understanding
Understanding urban scenes is a fundamental ability requirement for assisted driving and autonomous vehicles. Most of the available urban scene understanding methods use red-greenblue (RGB) images; however, their segme…
Autonomous VehiclesScene UnderstandingSegmentationThermal Image SegmentationRoadscapesQA: A Multitask, Multimodal Dataset for Visual Question Answering on Indian Roads
Understanding road scenes is essential for autonomous driving, as it enables systems to interpret visual surroundings to aid in effective decision-making. We present Roadscapes, a multitask multimodal dataset consisting …
Visual Question AnsweringScene UnderstandingAutonomous DrivingEfficient Multi-Task RGB-D Scene Analysis for Indoor Environments
Semantic scene understanding is essential for mobile agents acting in various environments. Although semantic segmentation already provides a lot of information, details about individual objects as well as the general sc…
Instance SegmentationPanoptic SegmentationScene ClassificationScene Classification (unified classes)+3Dynamic Resilient Spatio-Semantic Memory with Hybrid Localization for Mobile Manipulation
Reliable mobile manipulation in dynamic indoor environments requires a scene representation that remains geometrically consistent, semantically queryable, and computationally bounded as the environment changes. Existing …