paper-with-me

홈 › Papers

SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

2024-01-17 · Baoxiong Jia, Yixin Chen, Huangyue Yu, Yan Wang, Xuesong Niu, Tengyu Liu, Qing Li, Siyuan Huang

3D vision-language grounding, which focuses on aligning language with the 3D physical environment, stands as a cornerstone in the development of embodied agents. In comparison to recent advancements in the 2D domain, grounding language in 3D scenes faces several significant challenges: (i) the inherent complexity of 3D scenes due to the diverse object configurations, their rich attributes, and intricate relationships; (ii) the scarcity of paired 3D vision-language data to support grounded learning; and (iii) the absence of a unified learning framework to distill knowledge from grounded 3D data. In this work, we aim to address these three major challenges in 3D vision-language by examining the potential of systematically upscaling 3D vision-language learning in indoor environments. We introduce the first million-scale 3D vision-language dataset, SceneVerse, encompassing about 68K 3D indoor scenes and comprising 2.5M vision-language pairs derived from both human annotations and our scalable scene-graph-based generation approach. We demonstrate that this scaling allows for a unified pre-training framework, Grounded Pre-training for Scenes (GPS), for 3D vision-language learning. Through extensive experiments, we showcase the effectiveness of GPS by achieving state-of-the-art performance on all existing 3D visual grounding benchmarks. The vast potential of SceneVerse and GPS is unveiled through zero-shot transfer experiments in the challenging 3D vision-language tasks. Project website: https://scene-verse.github.io.

📄 PDF Abstract BibTeX arXiv:2401.09340

Code (0)

등록된 구현이 없습니다.

Tasks

3D visual groundingScene UnderstandingVisual Grounding

Methods 이 논문이 사용한 방법론

GPS Greedy Policy Search (GPS) is a simple algorithm that learns a policy for test-time data augmentation based on the predictive performance on a validation set. GPS starts with…

Similar Papers 제목 키워드 기반

Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding

2026-02-03 · Byeongju Woo, Zilin Wang, Byeonghyun Pak, Sangwoo Mo 외 arxiv

Vision-language models such as CLIP often struggle to faithfully understand long, detail-rich captions, relying on dominant scene cues while overlooking fine-grained visual evidence. We propose a hierarchical vision-lang…

Text Retrieval

Grounded 3D-LLM with Referent Tokens

2024-05-16 · Yilun Chen, Shuai Yang, Haifeng Huang, Tai Wang 외

Prior studies on 3D scene understanding have primarily developed specialized models for specific tasks or required task-specific fine-tuning. In this study, we propose Grounded 3D-LLM, which explores the potential of 3D …

Dense CaptioningDiversityInstruction FollowingLanguage Modeling+5

SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation

2026-03-24 · Ruisen Tu, Arth Shukla, Sohyun Yoo, Xuanlin Li 외 arxiv

Vision-Language-Action (VLA) models show promise for robotic control, yet performance in complex household environments remains sub-optimal. Mobile manipulation requires reasoning about global scene layout, fine-grained …

Does Structural Attention Improve Compositional Representations in Vision-Language Models?

2022-12-03 · NeurIPS Workshop: Self-Supervised Learning - Theory and Practice 2022 12 · Rohan Pandey, Rulin Shao, Paul Pu Liang, Louis-Philippe Morency

Although scaling self-supervised approaches has gained widespread success in Vision-Language pre-training, a number of works providing structural knowledge of visually-grounded semantics have recently shown incremental…

Visual Reasoning

Feature Splatting: Language-Driven Physics-Based Scene Synthesis and Editing

2024-04-01 · Ri-Zhao Qiu, Ge Yang, Weijia Zeng, Xiaolong Wang

Scene representations using 3D Gaussian primitives have produced excellent results in modeling the appearance of static and dynamic 3D scenes. Many graphics applications, however, demand the ability to manipulate both th…

Feature Splatting