paper-with-me

홈 › Papers

Mean-Flow based One-Step Vision-Language-Action

2026-03-02 · Yang Chen, Xiaoguang Ma, Bin Zhao arxiv

Recent advances in FlowMatching-based Vision-Language-Action (VLA) frameworks have demonstrated remarkable advantages in generating high-frequency action chunks, particularly for highly dexterous robotic manipulation tasks. Despite these notable achievements, their practical applications are constrained by prolonged generation latency, which stems from inherent iterative sampling requirements and architectural limitations. To address this critical bottleneck, we propose a Mean-Flow based One-Step VLA approach. Specifically, we resolve the noise-induced issues in the action generation process, thereby eliminating the consistency constraints inherent to conventional Flow-Matching methods. This significantly enhances generation efficiency and enables one-step action generation. Real-world robotic experiments show that the generation speed of the proposed Mean-Flow based One-Step VLA is 8.7 times and 83.9 times faster than that of SmolVLA and Diffusion Policy, respectively. These results elucidate its great potential as a high-efficiency backbone for VLA-based robotic manipulation.

📄 PDF Abstract BibTeX arXiv:2603.01469

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HybridFlow: A Two-Step Generative Policy for Robotic Manipulation

2026-02-14 · Zhenchen Dong, Jinna Fu, Jiaming Wu, Shengyuan Yu 외 arxiv

Limited by inference latency, existing robot manipulation policies lack sufficient real-time interaction capability with the environment. Although faster generation methods such as flow matching are gradually replacing d…

Robot ManipulationImage Generation

ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models

2026-03-18 · Zhou Fang, Jiaqi Wang, Yi Zhou, Qiongfeng Shi arxiv

Recent Vision-Language-Action (VLA) models equipped with Flow Matching (FM) action heads achieve state-of-the-art performance in complex robot manipulation. However, the multi-step iterative ODE solving required by FM in…

Robot Manipulation

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models

2026-05-21 · Zhi Liu arxiv

Vision-Language-Action (VLA) models have rapidly converged on a small set of architectural patterns: discrete-token autoregression (e.g. OpenVLA) and continuous-action flow-matching (e.g. pi-0.5). Yet preference alignmen…

OM2P: Offline Multi-Agent Mean-Flow Policy

2025-08-08 · Zhuoran Li, Xun Wang, Hai Zhong, Qingxin Xia 외 arxiv

Generative models, especially diffusion and flow-based models, have been promising in offline multi-agent reinforcement learning. However, integrating powerful generative models into this framework poses unique challenge…

Multi-agent Reinforcement Learning

LaMP: Learning Vision-Language-Action Policy with 3D Scene Flow as Latent Motion Prior

2026-03-26 · Xinkai Wang, Chenyi Wang, Yifu Xu, Mingzhe Ye 외 arxiv

We introduce \textbf{LaMP}, a dual-expert Vision-Language-Action framework that embeds dense 3D scene flow as a latent motion prior for robotic manipulation.Existing VLA models regress actions directly from 2D semantic v…