paper-with-me

Papers

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

2025-10-31 · John Won, Kyungmin Lee, Huiwon Jang, Dongyoung Kim, Jinwoo Shin arxiv

Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream diffusion (DUST), a world-model augmented VLA framework featuring a multimodal diffusion transformer that maintains separate modality streams while enabling cross-modal knowledge sharing. In addition, DUST utilizes independent noise perturbations and a decoupled flow matching loss to learn cross-modal causal relationships. We further introduce an asynchronous sampling method for action and vision tokens that enhances performance through inference-time scaling. Experimental results on simulated benchmarks like RoboCasa and GR-1 show that DUST achieves up to 6% gains over state-of-the-art VLA and world-modeling baselines, with inference-time scaling providing an additional 2-5% improvement. In real-world tasks using the Franka Research 3, DUST outperforms baselines by 10% in success rate. Finally, we demonstrate that DUST enables effective transfer learning through both pretraining on action-free videos and joint-training with heterogeneous robot and human datasets.

📄 PDF Abstract BibTeX arXiv:2510.27607

Code (0)

등록된 구현이 없습니다.

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

2026-07-23 · Sicheng Mo, Yuheng Li, Ziyang Leng, Krishna Kumar Singh 외 hf

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines …

Video Generation

D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction

2023-11-23 · Bowen Fu, Gu Wang, Chenyangguang Zhang, Yan Di 외

Reconstructing hand-held objects from a single RGB image is a challenging task in computer vision. In contrast to prior works that utilize deterministic modeling paradigms, we employ a point cloud denoising diffusion mod…

DenoisingObjectObject Reconstruction

Streamlining Industrial Contract Management with Retrieval-Augmented LLMs

2025-11-18 · Kristi Topollai, Tolga Dimlioglu, Anna Choromanska, Simon Odie 외 arxiv

Contract management involves reviewing and negotiating provisions, individual clauses that define rights, obligations, and terms of agreement. During this process, revisions to provisions are proposed and iteratively ref…

Synthetic Data Generation

Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing

2024-10-16 · Mingce Guo, Jingxuan He, Shengeng Tang, Zhangye Wang 외

Text-driven video editing utilizing generative diffusion models has garnered significant attention due to their potential applications. However, existing approaches are constrained by the limited word embeddings provided…

Video EditingWord Embeddings

RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding

2026-05-19 · Hanqing Liu, Mingjie Liu, Luoping Cui, Endian Lin 외 arxiv

Conventional vision-language models (VLMs) struggle to interpret scenes captured under adverse conditions (e.g., low light, high dynamic range, or fast motion) because standard RGB images degrade in such environments. Ev…

Scene Understanding