Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
Recent advancements in 3D Large Language Models (LLMs) have demonstrated promising capabilities for 3D scene understanding. However, previous methods exhibit deficiencies in general referencing and grounding capabilities for intricate scene comprehension. In this paper, we introduce the use of object identifiers and object-centric representations to interact with scenes at the object level. Specifically, we decompose the input 3D scene into a set of object proposals, each assigned a unique identifier token, which enables efficient object referencing and grounding during user-assistant interactions. Given the scarcity of scene-language data, we model the scene embeddings as a sequence of explicit object-level embeddings, derived from semantic-rich 2D or 3D representations. By employing object identifiers, we transform diverse 3D scene-language tasks into a unified question-answering format, facilitating joint training without the need for additional task-specific heads. With minimal fine-tuning on all downstream tasks, our model significantly outperforms existing methods on benchmarks including ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D.
Code (2)
Tasks
3D Question Answering (3D-QA)AttributeObjectQuestion AnsweringScene UnderstandingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
Recent advancements in multi-modal large language models (MLLMs) have shown strong potential for 3D scene understanding. However, existing methods struggle with fine-grained object grounding and contextual reasoning, lim…
Scene UnderstandingSpatial ReasoningBridging Text and Video: A Universal Multimodal Transformer for Video-Audio Scene-Aware Dialog
Audio-Visual Scene-Aware Dialog (AVSD) is a task to generate responses when chatting about a given video, which is organized as a track of the 8th Dialog System Technology Challenge (DSTC8). To solve the task, we propose…
Dialogue GenerationMulti-Task LearningText GenerationEditable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents
Scene simulation in autonomous driving has gained significant attention because of its huge potential for generating customized data. However, existing editable scene simulation approaches face limitations in terms of us…
Autonomous DrivingLanguage ModelingLanguage ModellingLarge Language Model+1ChatSplat: 3D Conversational Gaussian Splatting
Humans naturally interact with their 3D surroundings using language, and modeling 3D language fields for scene understanding and interaction has gained growing interest. This paper introduces ChatSplat, a system that con…
Large Language ModelScene UnderstandingChat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes
3D scene understanding has gained significant attention due to its wide range of applications. However, existing methods for 3D scene understanding are limited to specific downstream tasks, which hinders their practicali…
Language ModelingLanguage ModellingLarge Language ModelScene Understanding+1