paper-with-me

홈 › Papers

RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation

2024-12-11 · CVPR 2025 1 · Mingfei Han, Liang Ma, Kamila Zhumakhanova, Ekaterina Radionova, Jingyi Zhang, Xiaojun Chang, Xiaodan Liang, Ivan Laptev

Vision-and-Language Navigation (VLN) suffers from the limited diversity and scale of training data, primarily constrained by the manual curation of existing simulators. To address this, we introduce RoomTour3D, a video-instruction dataset derived from web-based room tour videos that capture real-world indoor spaces and human walking demonstrations. Unlike existing VLN datasets, RoomTour3D leverages the scale and diversity of online videos to generate open-ended human walking trajectories and open-world navigable instructions. To compensate for the lack of navigation data in online videos, we perform 3D reconstruction and obtain 3D trajectories of walking paths augmented with additional information on the room types, object locations and 3D shape of surrounding scenes. Our dataset includes $\sim$100K open-ended description-enriched trajectories with $\sim$200K instructions, and 17K action-enriched trajectories from 1847 room tour environments. We demonstrate experimentally that RoomTour3D enables significant improvements across multiple VLN tasks including CVDN, SOON, R2R, and REVERIE. Moreover, RoomTour3D facilitates the development of trainable zero-shot VLN agents, showcasing the potential and challenges of advancing towards open-world navigation.

📄 PDF Abstract BibTeX arXiv:2412.08591

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionDiversityVision and Language Navigation

Similar Papers 제목 키워드 기반

3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds

2025-12-28 · Ryousuke Yamada, Kohsuke Ide, Yoshihiro Fukuhara, Hirokatsu Kataoka 외 arxiv

Despite recent progress in 3D self-supervised learning, collecting large-scale 3D scene scans remains expensive and labor-intensive. In this work, we investigate whether 3D representations can be learned from unlabeled v…

Self-Supervised LearningRepresentation LearningInstance SegmentationPoint Clouds

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation

2025-12-28 · Kai Liu, Jungang Li, Yuchong Sun, Shengqiong Wu 외 arxiv

This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation. JavisGPT has a concise encoder-LLM-decoder architecture, which has a SyncFu…

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency

2025-11-30 · Zhongbin Guo, Jiahe Liu, Wenyu Gao, Yushan Li 외 arxiv

Text-driven 3D reconstruction requires masks that understand free-form instructions and remain stable under viewpoint changes. We present LISA-3D, a two-stage framework that adapts the instruction-following segmenter LIS…

Image Segmentation3D Reconstruction

Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation

2026-06-23 · Tianyu Zhu, Yingping Liang, Hesong Li, Ying Fu arxiv

Text-driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing models are typically trained on 2D image or video datasets with naive s…

Referring Video Object SegmentationZero-shot GeneralizationImage SegmentationVideo Segmentation

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

2025-05-26 · Zhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang 외

The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes, aiming for human-like visual-spatial intelligence. Nevertheless, achieving de…

3D ReconstructionSpatial Reasoning