paper-with-me

홈 › Papers

MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation

2026-03-26 · Yang Liu, Pengxiang Ding, Tengyue Jiang, Xudong Wang, Wenxuan Song, Minghui Lin, Han Zhao, Hongyin Zhang, Zifeng Zhuang, Wei Zhao, Siteng Huang, Jinkui Shi, Donglin Wang arxiv

Vision-Language-Action (VLA) models aim to control robots for manipulation from visual observations and natural-language instructions. However, existing hierarchical and autoregressive paradigms often introduce architectural overhead, suffer from temporal inconsistency and long-horizon error accumulation, and lack a mechanism to capture environment dynamics without extra modules. To this end, we present MMaDA-VLA, a fully native pre-trained large diffusion VLA model that unifies multi-modal understanding and generation in a single framework. Our key idea is a native discrete diffusion formulation that embeds language, images, and continuous robot controls into one discrete token space and trains a single backbone with masked token denoising to jointly generate a future goal observation and an action chunk in parallel. Iterative denoising enables global, order-free refinement, improving long-horizon consistency while grounding actions in predicted future visual outcomes without auxiliary world models. Experiments across simulation benchmarks and real-world tasks show state-of-the-art performance, achieving 98.0% average success on LIBERO and 4.78 average length on CALVIN.

📄 PDF Abstract BibTeX arXiv:2603.25406

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MMaDA: Multimodal Large Diffusion Language Models

2025-05-21 · Ling Yang, Ye Tian, Bowen Li, Xinchen Zhang 외

We introduce MMaDA, a novel class of multimodal diffusion foundation models designed to achieve superior performance across diverse domains such as textual reasoning, multimodal understanding, and text-to-image generatio…

Image GenerationReinforcement Learning (RL)Text to Image GenerationText-to-Image Generation

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

2025-11-12 · Ye Tian, Ling Yang, Jiongfan Yang, Anran Wang 외 arxiv

While thinking-aware generation aims to improve performance on complex tasks, we identify a critical failure mode where existing sequential, autoregressive approaches can paradoxically degrade performance due to error pr…

Reinforcement Learning

Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space

2026-01-19 · Zihang Wang, Siyue Zhang, Yilun Zhao, Jingyi Yang 외 arxiv

Embedding models are a fundamental component of modern AI systems such as semantic search and retrieval-augmented generation. Recent advances in large foundation models have substantially accelerated the development of e…

Visual Question AnsweringInformation Retrieval

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models

2025-12-22 · Yi Xin, Siqi Luo, Tianxiang Xu, Qi Qin 외 arxiv

Diffusion Multi-modal Large Language Models (dMLLMs) have recently emerged as a novel architecture unifying image generation and understanding. However, developing effective and efficient Test-Time Scaling (TTS) methods …

Image Generation

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models

2026-05-16 · Siqi Luo, Jianghan Shen, Yi Xin, Huayu Zheng 외 arxiv

Diffusion Multi-Modal Large Language Models (dMLLMs) are powerful for image generation, but optimizing them through reinforcement learning (RL) remains a major challenge. One primary difficulty is that a single image can…

Hierarchical Reinforcement LearningImage Generation