paper-with-me

Papers

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly

2026-04-10 · Zhi Jing, Jinbin Qiao, Ouyang Lu, Jicong Ao, Shuang Qiu, Huazhe Xu, Yu-Gang Jiang, Chenjia Bai arxiv

Spatial reasoning is a fundamental capability for embodied intelligence, especially for fine-grained manipulation tasks such as robotic assembly. Recent methods based on vision-language models (VLMs) largely rely on coarse 2D perception and struggle to perform accurate reasoning over complex 3D geometry. To address this limitation, we propose AssemLM, a spatial multimodal large language model for robotic assembly that integrates assembly manuals, point clouds, and textual instructions to predict task-critical 6D assembly poses with explicit geometric understanding. To bridge raw 3D perception and high-level linguistic reasoning, AssemLM employs a specialized point cloud encoder to extract fine-grained geometric and rotational features for accurate 3D spatial reasoning in assembly tasks. In addition, we introduce AssemBench, a large-scale benchmark for assembly-oriented spatial reasoning with over 900K multimodal samples and precise 6D pose annotations, extending evaluation from 2D grounding to full 3D geometric inference. Extensive experiments and real-robot evaluations demonstrate that AssemLM achieves state-of-the-art 6D pose reasoning performance and effectively supports fine-grained, multi-step assembly tasks in real-world settings. Code, models, and the AssemBench dataset will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2604.08983

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningPoint Clouds

Similar Papers 제목 키워드 기반

Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks

2025-10-29 · Xu Zheng, Zihao Dongfang, Lutao Jiang, Boyuan Zheng 외 arxiv

Humans possess spatial reasoning abilities that enable them to understand spaces through multimodal observations, such as vision and sound. Large multimodal reasoning models extend these abilities by learning to perceive…

Vision-Language NavigationVisual Question AnsweringMultimodal ReasoningSpatial Reasoning

Open3DVQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space

2025-03-14 · Weichen Zhang, Zile Zhou, Zhiheng Zheng, Chen Gao 외

Spatial reasoning is a fundamental capability of embodied agents and has garnered widespread attention in the field of multimodal large language models (MLLMs). In this work, we propose a novel benchmark, Open3DVQA, to c…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+2

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning

2026-04-19 · Yian Li, Yang Jiao, Bin Zhu, Tianwen Qian 외 arxiv

Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language models. Despite promising performance, r…

multimodal generationSpatial Reasoning

An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

2024-11-09 · Fatemeh Shiri, Xiao-Yu Guo, Mona Golestan Far, Xin Yu 외

Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. In this paper, we construct a novel VQA d…

object-detectionObject DetectionSpatial ReasoningVisual Question Answering (VQA)

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning

2026-03-24 · Jiacheng Hua, Yishu Yin, Yuhang Wu, Tai Wang 외 arxiv

Existing Multimodal Large Language Models (MLLMs) struggle with 3D spatial reasoning, as they fail to construct structured abstractions of the 3D environment depicted in video inputs. To bridge this gap, drawing inspirat…

Question AnsweringSpatial Reasoning