paper-with-me

Papers

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models

2025-07-22 · Xiaoyan Wang, Zeju Li, Yifan Xu, Jiaxing Qi, Zhifei Yang, Ruifei Ma, Xiangde Liu, Chao Zhang arxiv

New era has unlocked exciting possibilities for extending Large Language Models (LLMs) to tackle 3D vision-language tasks. However, most existing 3D multimodal LLMs (MLLMs) rely on compressing holistic 3D scene information or segmenting independent objects to perform these tasks, which limits their spatial awareness due to insufficient representation of the richness inherent in 3D scenes. To overcome these limitations, we propose Spatial 3D-LLM, a 3D MLLM specifically designed to enhance spatial awareness for 3D vision-language tasks by enriching the spatial embeddings of 3D scenes. Spatial 3D-LLM integrates an LLM backbone with a progressive spatial awareness scheme that progressively captures spatial information as the perception field expands, generating location-enriched 3D scene embeddings to serve as visual prompts. Furthermore, we introduce two novel tasks: 3D object distance measurement and 3D layout editing, and construct a 3D instruction dataset, MODEL, to evaluate the model's spatial awareness capabilities. Experimental results demonstrate that Spatial 3D-LLM achieves state-of-the-art performance across a wide range of 3D vision-language tasks, revealing the improvements stemmed from our progressive spatial awareness scheme of mining more profound spatial information. Our code is available at https://github.com/bjshuyuan/Spatial-3D-LLM.

📄 PDF Abstract BibTeX arXiv:2507.16524

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation

2026-03-10 · Jiahang Liu, Tianyu Xu, Jiawei Chen, Lu Yue 외 arxiv

Recent embodied navigation approaches leveraging Vision-Language Models (VLMs) demonstrate strong generalization in versatile Vision-Language Navigation (VLN). However, reliable path planning in complex environments rema…

Vision-Language Navigation

Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

2024-04-11 · CVPR 2024 1 · Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo 외

Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly for visual question answering (VQA). How…

DescriptiveHallucinationQuestion AnsweringSpatial Reasoning+4

Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models

2025-03-21 · Jianing Qi, Jiawei Liu, Hao Tang, Zhigang Zhu

Vision-Language Models (VLMs) excel at identifying and describing objects but struggle with spatial reasoning such as accurately understanding the relative positions of objects. Inspired by the dual-pathway (ventral-dors…

DiagnosticObject RecognitionSpatial Reasoning

The Spatial Blindspot of Vision-Language Models

2026-01-15 · Nahid Alam, Leema Krishna Murali, Siddhant Bharadwaj, Patrick Liu 외 arxiv

Vision-language models (VLMs) have advanced rapidly, but their ability to capture spatial relationships remains a blindspot. Current VLMs are typically built with contrastive language-image pretraining (CLIP) style image…

Spatial Reasoning

Enhancing the Spatial Awareness Capability of Multi-Modal Large Language Model

2023-10-31 · Yongqiang Zhao, Zhenyu Li, Zhi Jin, Feng Zhang 외

The Multi-Modal Large Language Model (MLLM) refers to an extension of the Large Language Model (LLM) equipped with the capability to receive and infer multi-modal data. Spatial awareness stands as one of the crucial abil…

Autonomous DrivingLanguage ModelingLanguage ModellingLarge Language Model+2