paper-with-me

홈 › Papers

Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

2025-04-02 · Haochen Wang, Yucheng Zhao, Tiancai Wang, Haoqiang Fan, Xiangyu Zhang, Zhaoxiang Zhang

The rapid development of Large Multimodal Models (LMMs) for 2D images and videos has spurred efforts to adapt these models for interpreting 3D scenes. However, the absence of large-scale 3D vision-language datasets has posed a significant obstacle. To address this issue, typical approaches focus on injecting 3D awareness into 2D LMMs by designing 3D input-level scene representations. This work provides a new perspective. We introduce reconstructive visual instruction tuning with 3D-awareness (Ross3D), which integrates 3D-aware visual supervision into the training procedure. Specifically, it incorporates cross-view and global-view reconstruction. The former requires reconstructing masked views by aggregating overlapping information from other views. The latter aims to aggregate information from all available views to recover Bird's-Eye-View images, contributing to a comprehensive overview of the entire scene. Empirically, Ross3D achieves state-of-the-art performance across various 3D scene understanding benchmarks. More importantly, our semi-supervised experiments demonstrate significant potential in leveraging large amounts of unlabeled 3D vision-only data.

📄 PDF Abstract BibTeX arXiv:2504.01901

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Reconstructive Visual Instruction Tuning

2024-10-12 · Haochen Wang, Anlin Zheng, Yucheng Zhao, Tiancai Wang 외

This paper introduces reconstructive visual instruction tuning (ROSS), a family of Large Multimodal Models (LMMs) that exploit vision-centric supervision signals. In contrast to conventional visual instruction tuning app…

Denoising

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

2025-05-26 · Zhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang 외

The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes, aiming for human-like visual-spatial intelligence. Nevertheless, achieving de…

3D ReconstructionSpatial Reasoning

Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

2024-04-11 · CVPR 2024 1 · Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo 외

Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly for visual question answering (VQA). How…

DescriptiveHallucinationQuestion AnsweringSpatial Reasoning+4

On the Loss of Context-awareness in General Instruction Fine-tuning

2024-11-05 · Yihan Wang, Andrew Bai, Nanyun Peng, Cho-Jui Hsieh

Pre-trained Large Language Models (LLMs) require post-training methods such as supervised fine-tuning (SFT) on instruction-response pairs to enable instruction following. However, this process can potentially harm existi…

BenchmarkingInstruction Following

Reg3D: Reconstructive Geometry Instruction Tuning for 3D Scene Understanding

2025-09-03 · Hongpei Zheng, Lintao Xiang, Qijun Yang, Qian Lin 외 arxiv

The rapid development of Large Multimodal Models (LMMs) has led to remarkable progress in 2D visual understanding; however, extending these capabilities to 3D scene understanding remains a significant challenge. Existing…

Scene UnderstandingSpatial Reasoning