paper-with-me

홈 › Papers

Mono-hydra: Real-time 3D scene graph construction from monocular camera input with IMU

2023-08-10 · U. V. B. L. Udugama, G. Vosselman, F. Nex

The ability of robots to autonomously navigate through 3D environments depends on their comprehension of spatial concepts, ranging from low-level geometry to high-level semantics, such as objects, places, and buildings. To enable such comprehension, 3D scene graphs have emerged as a robust tool for representing the environment as a layered graph of concepts and their relationships. However, building these representations using monocular vision systems in real-time remains a difficult task that has not been explored in depth. This paper puts forth a real-time spatial perception system Mono-Hydra, combining a monocular camera and an IMU sensor setup, focusing on indoor scenarios. However, the proposed approach is adaptable to outdoor applications, offering flexibility in its potential uses. The system employs a suite of deep learning algorithms to derive depth and semantics. It uses a robocentric visual-inertial odometry (VIO) algorithm based on square-root information, thereby ensuring consistent visual odometry with an IMU and a monocular camera. This system achieves sub-20 cm error in real-time processing at 15 fps, enabling real-time 3D scene graph construction using a laptop GPU (NVIDIA 3080). This enhances decision-making efficiency and effectiveness in simple camera setups, augmenting robotic system agility. We make Mono-Hydra publicly available at: https://github.com/UAV-Centre-ITC/Mono_Hydra

📄 PDF Abstract BibTeX arXiv:2308.05515

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingGPUgraph constructionNavigateVisual Odometry

Similar Papers 제목 키워드 기반

Mono-Hydra++: Real-Time Monocular Scene Graph Construction with Multi-Task Learning for 3D Indoor Mapping

2026-05-17 · U. V. B. L. Udugama, George Vosselman, Francesco Nex arxiv

Autonomous agile robots need more than metric geometry: they must understand objects, rooms, places, and spatial relations for search, inspection, exploration, and human robot interaction. Conventional metric maps suppor…

Multi-Task LearningCollision Avoidance

Hydra++: Real-Time Hierarchical 3D Scene Graph Construction With Object-Level Shape Estimation

2026-07-10 · Hyungtae Lim, Nathan Hughes, Xihang Yu, Ruihan Xu 외 arxiv

3D scene graphs provide a hierarchical abstraction of environments by encoding spatial entities, such as objects and places, and their relationships. However, existing scene graph systems model object geometry coarsely, …

Point Clouds

M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception

2025-10-20 · U. V. B. L Udugama, George Vosselman, Francesco Nex arxiv

Deploying real-time spatial perception on edge devices requires efficient multi-task models that leverage complementary task information while minimizing computational overhead. This paper introduces Multi-Mono-Hydra (M2…

Computational EfficiencySemantic SegmentationMulti-Task Learning

PhyDrawGen: Physically Grounded Diagram Generation from Natural Language

2026-05-28 · Nafiul Haque, Syed Nazmus Sakib, Shifat E Arman arxiv

Generating physics diagrams from text requires strict adherence to physical laws. While current generative models produce visually plausible outputs, they systematically hallucinate force vectors, ignore conservation law…

Scene Understanding

Graph neural network-based surrogate modelling for real-time hydraulic prediction of urban drainage networks

2024-04-16 · ZhiYu Zhang, Chenkaixiang Lu, Wenchong Tian, Zhenliang Liao 외

Physics-based models are computationally time-consuming and infeasible for real-time scenarios of urban drainage networks, and a surrogate model is needed to accelerate the online predictive modelling. Fully-connected ne…

Graph Neural NetworkPrediction