paper-with-me

Papers

From Geometric Labels to Semantic Understanding of Indoor Building Components Using Multimodal Large Language Models

2026-07-04 · Shuju Jing, Chao Yin arxiv

Point cloud-based understanding has become an important enabler for facility operation and maintenance involving indoor building components. However, existing methods output only discrete labels without explaining component functions or natural language interactions. This paper proposes Building-MLLM, a point cloud-centered multimodal large language model (MLLM) for indoor components, which models point clouds and instructions to generate responses across Simple Recognition, Complex Captioning, and Multi-Engineering Question Answering tasks. Building-MLLM addresses semantic concentration through four domain-specific mechanisms: Point Information Enhancer for task-relevant semantics, Geometry-Preserving Regularization preventing geometric erosion, fixed textual prefix for domain stabilization, and multi-dimensional LoRA balancing recognition with reasoning. A multi-constraint progressive instruction-generation engine is developed to compile a synthetic point cloud-text dataset with 4198 objects, 37,782 instruction-following pairs, and 47 categories. Experiments show that Building-MLLM achieves 88.00%, 65.10%, and 68.14% on the three task types, respectively, demonstrating superior indoor component language understanding and providing initial generalizability in transfer inference on other real-world datasets.

📄 PDF Abstract BibTeX arXiv:2607.03661

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringPoint Clouds

Similar Papers 제목 키워드 기반

Incremental Semantics-Aided Meshing from LiDAR-Inertial Odometry and RGB Direct Label Transfer

2026-04-10 · Muhammad Affan, Ville Lehtola, George Vosselman arxiv

Geometric high-fidelity mesh reconstruction from LiDAR-inertial scans remains challenging in large, complex indoor environments -- such as cultural buildings -- where point cloud sparsity, geometric drift, and fixed fusi…

3DSES: an indoor Lidar point cloud segmentation dataset with real and pseudo-labels from a 3D model

2025-01-29 · Maxime Mérizette, Nicolas Audebert, Pierre Kervella, Jérôme Verdun

Semantic segmentation of indoor point clouds has found various applications in the creation of digital twins for robotics, navigation and building information modeling (BIM). However, most existing datasets of labeled in…

Point Cloud SegmentationSemantic Segmentation

First Shape, Then Meaning: Efficient Geometry and Semantics Learning for Indoor Reconstruction

2026-05-05 · Remi Chierchia, Léo Lebrat, David Ahmedt-Aristizabal, Olivier Salvado 외 arxiv

Neural Surface Reconstruction has become a standard methodology for indoor 3D reconstruction, with Signed Distance Functions (SDFs) proving particularly effective for representing scene geometry. A variety of application…

3D Reconstruction

Understanding Indoor Scenes Using 3D Geometric Phrases

2013-06-01 · CVPR 2013 6 · Wongun Choi, Yu-Wei Chao, Caroline Pantofaru, Silvio Savarese

Visual scene understanding is a difficult problem interleaving object detection, geometric reasoning and scene classification. We present a hierarchical scene model for learning and reasoning about complex indoor scenes …

General ClassificationObjectobject-detectionObject Detection+3

Intelligent Spatial Perception by Building Hierarchical 3D Scene Graphs for Indoor Scenarios with the Help of LLMs

2025-03-19 · Yao Cheng, Zhe Han, Fengyang Jiang, Huaizhen Wang 외

This paper addresses the high demand in advanced intelligent robot navigation for a more holistic understanding of spatial environments, by introducing a novel system that harnesses the capabilities of Large Language Mod…

ObjectRobot NavigationTask Planning