paper-with-me

홈 › Papers

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

2026-03-31 · Yi Chen, Yuying Ge, Hui Zhou, Mingyu Ding, Yixiao Ge, Xihui Liu arxiv

The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). However, most existing end-to-end VLAs treat the VLM primarily as a multimodal encoder, directly mapping vision-language features to low-level actions. This paradigm underutilizes the VLM's potential in high-level decision making and introduces training instability, frequently degrading its rich semantic representations. To address these limitations, we introduce DIAL, a framework bridging high-level decision making and low-level motor execution through a differentiable latent intent bottleneck. Specifically, a VLM-based System-2 performs latent world modeling by synthesizing latent visual foresight within the VLM's native feature space; this foresight explicitly encodes intent and serves as the structural bottleneck. A lightweight System-1 policy then decodes this predicted intent together with the current observation into precise robot actions via latent inverse dynamics. To ensure optimization stability, we employ a two-stage training paradigm: a decoupled warmup phase where System-2 learns to predict latent futures while System-1 learns motor control under ground-truth future guidance within a unified feature space, followed by seamless end-to-end joint optimization. This enables action-aware gradients to refine the VLM backbone in a controlled manner, preserving pre-trained knowledge. Extensive experiments on the RoboCasa GR1 Tabletop benchmark show that DIAL establishes a new state-of-the-art, achieving superior performance with 10x fewer demonstrations than prior methods. Furthermore, by leveraging heterogeneous human demonstrations, DIAL learns physically grounded manipulation priors and exhibits robust zero-shot generalization to unseen objects and novel configurations during real-world deployment on a humanoid robot.

📄 PDF Abstract BibTeX arXiv:2603.29844

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationDecision Making

Similar Papers 제목 키워드 기반

Latent Intention Dialogue Models

2017-05-29 · ICML 2017 8 · Tsung-Hsien Wen, Yishu Miao, Phil Blunsom, Steve Young

Developing a dialogue agent that is capable of making autonomous decisions and communicating by natural language is one of the long-term goals of machine learning research. Traditional approaches either rely on hand-craf…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Variational Inference

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos

2026-06-17 · Runze Xu, Yiluo Zhang, Jian Wang, Yu Wang 외 arxiv

Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets with high-fidelity action annotations. While egocentric human manipulation videos are abundant and capture signi…

LaDA: Latent Dialogue Action For Zero-shot Cross-lingual Neural Network Language Modeling

2023-08-05 · Zhanyu Ma, Jian Ye, Shuang Cheng

Cross-lingual adaptation has proven effective in spoken language understanding (SLU) systems with limited resources. Existing methods are frequently unsatisfactory for intent detection and slot filling, particularly for …

Intent DetectionLanguage ModelingLanguage Modellingslot-filling+2

Learn to Discover Dialog Intents via Self-supervised Context Pretraining

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Intent detection is one of most critical tasks in prevalent task-oriented dialog systems. However, most systems could only identify a fixed set of intents, without covering a ubiquitous space of real-world semantics. Ind…

Contrastive LearningIntent Detection

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue

2025-09-18 · Xingyao Lin, Xinghao Zhu, Tianyi Lu, Guojin Zhong 외 arxiv

Embodied agents are intelligent systems designed to perceive, reason, and act within the physical world. While the robotics community has long strived to build such versatile agents, a fundamental limitation persists: mo…