paper-with-me

홈 › Papers

World-Consistent Data Generation for Vision-and-Language Navigation

2024-12-09 · Yu Zhong, Rui Zhang, Zihao Zhang, Shuo Wang, Chuan Fang, Xishan Zhang, Jiaming Guo, Shaohui Peng, Di Huang, Yanyang Yan, Xing Hu, Ping Tan, Qi Guo

Vision-and-Language Navigation (VLN) is a challenging task that requires an agent to navigate through photorealistic environments following natural-language instructions. One main obstacle existing in VLN is data scarcity, leading to poor generalization performance over unseen environments. Tough data argumentation is a promising way for scaling up the dataset, how to generate VLN data both diverse and world-consistent remains problematic. To cope with this issue, we propose the world-consistent data generation (WCGEN), an efficacious data-augmentation framework satisfying both diversity and world-consistency, targeting at enhancing the generalizations of agents to novel environments. Roughly, our framework consists of two stages, the trajectory stage which leverages a point-cloud based technique to ensure spatial coherency among viewpoints, and the viewpoint stage which adopts a novel angle synthesis method to guarantee spatial and wraparound consistency within the entire observation. By accurately predicting viewpoint changes with 3D knowledge, our approach maintains the world-consistency during the generation procedure. Experiments on a wide range of datasets verify the effectiveness of our method, demonstrating that our data augmentation strategy enables agents to achieve new state-of-the-art results on all navigation tasks, and is capable of enhancing the VLN agents' generalization ability to unseen environments.

📄 PDF Abstract BibTeX arXiv:2412.06413

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationNavigateVision and Language Navigation

Similar Papers 제목 키워드 기반

Emu3.5: Native Multimodal Models are World Learners

2025-10-30 · Yufeng Cui, Honghao Chen, Haoge Deng, Xu Huang 외 arxiv

We introduce Emu3.5, a large-scale multimodal world model that natively predicts the next state across vision and language. Emu3.5 is pre-trained end-to-end with a unified next-token prediction objective on a corpus of v…

Reinforcement LearningMultimodal ReasoningImage Generation

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

2025-10-01 · Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li 외 arxiv

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Visi…

Trajectory ForecastingTrajectory PredictionVideo Generation

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

2026-07-03 · Zelin Zhao, Min Shi, Bo Yuan, Haotian Xue 외 arxiv

World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-World (VLAW) modeling. Meanwhile, unified vi…

multimodal generation

World Action Models: A Survey

2026-06-18 · Qiuhong Shen, Shihua Zhang, Yue Liao, Qi Li 외 arxiv

World Action Models (WAMs) are embodied predictive-action models that make a forecast of the future available to action. Recent WAMs repurpose large video generation models, and a parallel line relies on language or visi…

Video Generation

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models

2026-05-11 · Zuojin Tang, Haoyun Liu, Xinyuan Chang, Changjie Wu 외 arxiv

Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evidence of how the physical world changes. Latent action models offer a pr…