paper-with-me

홈 › Papers

Unified Visuomotor Targets: Supervising VLAs Beyond Physical Actions

2026-08-04 · Zhenyang Feng, Unnat Jain arxiv

VLA models are trained to predict robot actions from visual and language observations. This is a natural choice, but it creates a mismatch: VLMs encode rich, high-level representations of scenes and goals, while robot actions are low-level signals with limited task structure. We ask whether changing what the policy is trained to predict, rather than how it is architecturally designed, can yield better and more efficiently trained policies. We propose UVT (Unified Visuomotor Target), a unified latent prediction target that jointly encodes motor control and visual scene transition information, requiring no architectural changes and no additional data. Applied to two representative VLA systems across simulation benchmarks and real bimanual manipulation tasks, UVT improves training efficiency, final task performance, and policy robustness, with particularly strong gains under limited training budgets and challenging environmental conditions. Rollout videos and additional qualitative results are available at our project webpage: https://unified-visuomotor-targets.github.io/

📄 PDF Abstract BibTeX arXiv:2608.03563

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model

2025-11-03 · Wenqi Liang, Gan Sun, Yao He, Jiahua Dong 외 arxiv

Vision-Language-Action models (VLAs) are emerging as powerful tools for learning generalizable visuomotor control policies. However, current VLAs are mostly trained on large-scale image-text-action data and remain limite…

Scene Understanding

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models

2026-05-06 · Yihan Lin, Haoyang Li, Yang Li, Haitao Shen 외 arxiv

Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions ar…

Grounded World Model for Semantically Generalizable Planning

2026-04-13 · Quanyi Li, Lan Feng, Haonan Zhang, Wuyang Li 외 arxiv

In Model Predictive Control (MPC), world models predict the future outcomes of various action proposals, which are then scored to guide the selection of the optimal action. For visuomotor MPC, the score function is a dis…

RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models

2025-06-21 · Jacky Kwok, Christopher Agia, Rohan Sinha, Matt Foutter 외

Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in visuomotor control, yet ensuring their robustness in unstructured real-world environments remains a persistent challenge. In this paper, we…

Synthetic Data GenerationVision-Language-Action

Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

2026-05-21 · Pengteng Li, Weiyu Guo, He Zhang, Tiefu Cai 외 arxiv

We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly assume that task-relevant objects are always visible, leading to brittl…