paper-with-me

홈 › Papers

From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges

2026-04-23 · Yiming Zhong, Yaoyu He, Zemin Yang, Pengfei Tian, Yifan Huang, Qingqiu Huang, Xinge Zhu, Yuexin Ma arxiv

Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal scale mismatch between cognition and action. Existing generative VLA policies typically adopt a "Generation-from-Noise" paradigm, which disregards this disparity, leading to representation inefficiency and weak condition alignment during optimization. In this work, we propose ResVLA, an architecture that shifts the paradigm to "Refinement-from-Intent." Recognizing that robotic motion naturally decomposes into global intent and local dynamics, ResVLA utilizes spectral analysis to decouple control into a deterministic low-frequency anchor and a stochastic high-frequency residual. By anchoring the generative process on the predicted intent, our model focuses strictly on refining local dynamics via a residual diffusion bridge. Extensive simulation experiments show that ResVLA achieves competitive performance, strong robustness to language and robot embodiment perturbations, and faster convergence than standard generative baselines. ResVLA also demonstrates strong performance in real-world robot experiments.

📄 PDF Abstract BibTeX arXiv:2604.21391

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fidelity-Constrained Anchoring for Black-Box Denoisers

2026-08-13 · Masaki Satoh arxiv

We propose a fidelity-constrained framework that anchors the output of a black-box denoiser to its input without retraining and with little additional computation. The method linearly blends the denoised image with the i…

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

2026-07-07 · Hanan Gani, Tejal Kulkarni, Madhoolika Chodavarapu, Nicklas Hansen 외 arxiv

Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficu…

One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow

2025-11-17 · Zeyuan Wang, Da Li, Yulin Chen, Ye Shi 외 arxiv

We introduce a one-step generative policy for offline reinforcement learning that maps noise directly to actions via a residual reformulation of MeanFlow, making it compatible with Q-learning. While one-step Gaussian pol…

Reinforcement Learning

RFS: Reinforcement Learning with Residual Flow Steering for Dexterous Manipulation

2026-02-02 · Entong Su, Tyler Westenbroek, Anusha Nagabandi, Abhishek Gupta arxiv

Imitation learning has emerged as an effective approach for bootstrapping sequential decision-making in robotics, achieving strong performance even in high-dimensional dexterous manipulation tasks. Recent behavior clonin…

Reinforcement Learning

SUREFlow: State-space Uncertainty-aware REsidual Flow Matching for Robust Robot Manipulation

2026-07-11 · Md Tanvir Islam, Sai Navaneet Peddapalli, Sangmoon Lee, Sangtae Ahn arxiv

Generative vision-language-action policies have advanced robot manipulation, but they often exhibit instability under noise, partial observability, and stochastic initial conditions. During extended rollouts, small veloc…

Computational EfficiencyRobot Manipulation