paper-with-me

홈 › Papers

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

2026-07-03 · Zelin Zhao, Min Shi, Bo Yuan, Haotian Xue, Jialuo Li, Lama Moukheiber, Humphrey Shi, Yongxin Chen arxiv

World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-World (VLAW) modeling. Meanwhile, unified vision-language models have demonstrated strong multimodal generation capabilities, yet their potential as world models remains underexplored. In this work, we introduce \texttt{WorldBagel}, a unified VLAW framework built on BAGEL, a modern multimodal unified model, and use it to systematically investigate the role of unification in world modeling. Across multi-task robotic manipulation and cross-domain experiments, \texttt{WorldBagel} consistently outperforms task-specific alternatives and learns action representations that are more structured and semantically aligned with visual and linguistic context. Experiments on LIBERO, Language Table, and Franka show that unification is not only an architectural convenience, but also a key factor in learning effective VLAW models, leading to consistent empirical gains and deeper insights into multimodal world modeling. Code and model checkpoints will be released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2607.03461

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal generation

Similar Papers 제목 키워드 기반

Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System

2026-09-01 · Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu 외 hf

While unified multimodal models (UMMs) jointly perform visual understanding and generation within a single model, functional unification does not guarantee learning synergy: the two objectives may reinforce each other, c…

PUMA: Empowering Unified MLLM with Multi-granular Visual Generation

2024-10-17 · Rongyao Fang, Chengqi Duan, Kun Wang, Hao Li 외

Recent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models (MLLMs) for vi…

DiversityImage GenerationImage ManipulationText to Image Generation+1

UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation

2025-05-15 · Yi Li, Haonan Wang, Qixiang Zhang, Boyu Xiao 외

The emergence of unified multimodal understanding and generation models is rapidly attracting attention because of their ability to enhance instruction-following capabilities while minimizing model redundancy. However, t…

DiversityInstruction Following

OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation

2025-09-03 · Han Li, Xinyu Peng, Yaoming Wang, Zelin Peng 외 arxiv

We introduce OneCAT, a unified multimodal model that seamlessly integrates understanding, generation, and editing within a novel, pure decoder-only transformer architecture. Our framework uniquely eliminates the need for…

multimodal generation

UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets

2025-09-18 · Pengyu Wang, Shaojun Zhou, Chenkun Tan, Xinghao Wang 외 arxiv

Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image s…

Visual Question Answering