paper-with-me

Papers

Towards Multimodal Multitask Scene Understanding Models for Indoor Mobile Agents

2022-09-27 · Yao-Hung Hubert Tsai, Hanlin Goh, Ali Farhadi, Jian Zhang

The perception system in personalized mobile agents requires developing indoor scene understanding models, which can understand 3D geometries, capture objectiveness, analyze human behaviors, etc. Nonetheless, this direction has not been well-explored in comparison with models for outdoor environments (e.g., the autonomous driving system that includes pedestrian prediction, car detection, traffic sign recognition, etc.). In this paper, we first discuss the main challenge: insufficient, or even no, labeled data for real-world indoor environments, and other challenges such as fusion between heterogeneous sources of information (e.g., RGB images and Lidar point clouds), modeling relationships between a diverse set of outputs (e.g., 3D object locations, depth estimation, and human poses), and computational efficiency. Then, we describe MMISM (Multi-modality input Multi-task output Indoor Scene understanding Model) to tackle the above challenges. MMISM considers RGB images as well as sparse Lidar points as inputs and 3D object detection, depth completion, human pose estimation, and semantic segmentation as output tasks. We show that MMISM performs on par or even better than single-task models; e.g., we improve the baseline 3D object detection results by 11.7% on the benchmark ARKitScenes dataset.

📄 PDF Abstract BibTeX arXiv:2209.13156

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAutonomous DrivingComputational EfficiencyDepth CompletionDepth EstimationObjectobject-detectionObject DetectionPose EstimationScene UnderstandingSemantic SegmentationTraffic Sign Recognition

Similar Papers 제목 키워드 기반

Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning

2025-06-06 · Sheng Chen, Peiyu He, Jiaxin Hu, Ziyang Liu 외

Modern robot navigation systems encounter difficulties in diverse and complex indoor environments. Traditional approaches rely on multiple modules with small models or rule-based systems and thus lack adaptability to new…

Robot NavigationSelf-Supervised LearningVisual Place Recognition

MTANet: Multitask-Aware Network With Hierarchical Multimodal Fusion for RGB-T Urban Scene Understanding

2022-04-05 · journal 2022 4 · WuJie Zhou, Shaohua Dong, Jingsheng Lei, Lu Yu

Understanding urban scenes is a fundamental ability requirement for assisted driving and autonomous vehicles. Most of the available urban scene understanding methods use red-greenblue (RGB) images; however, their segme…

Autonomous VehiclesScene UnderstandingSegmentationThermal Image Segmentation

RoadscapesQA: A Multitask, Multimodal Dataset for Visual Question Answering on Indian Roads

2026-02-13 · Vijayasri Iyer, Maahin Rathinagiriswaran, Jyothikamalesh S arxiv

Understanding road scenes is essential for autonomous driving, as it enables systems to interpret visual surroundings to aid in effective decision-making. We present Roadscapes, a multitask multimodal dataset consisting …

Visual Question AnsweringScene UnderstandingAutonomous Driving

Efficient Multi-Task RGB-D Scene Analysis for Indoor Environments

2022-07-10 · Daniel Seichter, Söhnke Benedikt Fischedick, Mona Köhler, Horst-Michael Groß

Semantic scene understanding is essential for mobile agents acting in various environments. Although semantic segmentation already provides a lot of information, details about individual objects as well as the general sc…

Instance SegmentationPanoptic SegmentationScene ClassificationScene Classification (unified classes)+3

Dynamic Resilient Spatio-Semantic Memory with Hybrid Localization for Mobile Manipulation

2026-05-30 · Zhijie Yan, Shufei Li, Ze Zhang, Xin Liu 외 arxiv

Reliable mobile manipulation in dynamic indoor environments requires a scene representation that remains geometrically consistent, semantically queryable, and computationally bounded as the environment changes. Existing …