paper-with-me

홈 › Papers

E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion

2025-11-26 · Zhihao Zhan, Jiaying Zhou, Likui Zhang, Qinhan Lv, Hao Liu, Jusheng Zhang, Weizheng Li, Ziliang Chen, Tianshui Chen, Ruifeng Zhai, Keze Wang, Liang Lin, Guangrun Wang arxiv

Vision-Language-Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, existing VLA systems still struggle to generalize across diverse tasks, scenes, and camera viewpoints, and often produce coarse or unstable actions. We argue that these limitations are closely tied to the structural properties of actions in VLA settings, including the inherent multi-peaked nature of action distributions, the token-based symbolic reasoning of pretrained VLM/VLA backbones, and the effective finite resolution imposed by real-world robotic control. Motivated by these properties, we introduce E0, a tweedie discrete diffusion framework that formulates action generation as iterative denoising over quantized action tokens. By operating in a discrete action space with a principled diffusion process, E0 naturally aligns with token-based reasoning, supports fine-grained yet executable action control, and avoids the distributional mismatch of masking-based discrete diffusion. We further introduce a spherical viewpoint perturbation augmentation to enhance robustness to camera shifts without additional data. Experiments on LIBERO, VLABench, ManiSkill, and a real-world Franka arm demonstrate that E0 achieves state-of-the-art performance across 14 diverse environments, outperforming strong baselines by 10.7% on average.

📄 PDF Abstract BibTeX arXiv:2511.21542

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching

2026-05-20 · Jangho Park, Geon Yeong Park, Gihyun Kwon, Jong Chul Ye arxiv

Extending the generation horizon of video diffusion models to long sequences remains a long-standing and important challenge. Existing training-free approaches fall into two categories: extensions of bidirectional models…

Video Generation

ViCTr: Vital Consistency Transfer for Pathology Aware Image Synthesis

2025-05-08 · Onkar Susladkar, Gayatri Deshmukh, Yalcin Tur, Ulas Bagci

Synthesizing medical images remains challenging due to limited annotated pathological data, modality domain gaps, and the complexity of representing diffuse pathologies such as liver cirrhosis. Existing methods often str…

8kData AugmentationImage Generation

Beyond First-Order Tweedie: Solving Inverse Problems using Latent Diffusion

2023-12-01 · CVPR 2024 1 · Litu Rout, Yujia Chen, Abhishek Kumar, Constantine Caramanis 외

Sampling from the posterior distribution poses a major computational challenge in solving inverse problems using latent diffusion models. Common methods rely on Tweedie's first-order moments, which are known to induce a …

text-guided-image-editing

Vessel Traffic Flow Prediction on Sparse Data via Spatio-Temporal Graph Neural Networks with a Learnable Tweedie Head

2026-06-05 · Kyeongjun Lee, Heeyoung Kim arxiv

Accurate vessel traffic flow prediction is crucial for smart port operations and navigational safety. However, maritime traffic flow data are often highly sparse with intermittent bursts, making robust forecasting challe…

The Bregman-Tweedie Classification Model

2019-07-16 · Hyenkyun Woo

This work proposes the Bregman-Tweedie classification model and analyzes the domain structure of the extended exponential function, an extension of the classic generalized exponential function with additional scaling par…

ClassificationGeneral Classificationmodel