paper-with-me

Papers

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

2026-09-11 · Jianman Lin, Shailesh Shailesh, Zhongyi Luo, Jiafei Duan hf

Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such vision-action shortcuts can undermine generalization when these correlations change under distribution shifts. Mitigating these shortcuts requires constraining how visual information is used for action generation while preserving task-relevant spatial information. We propose Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface. Stage 1 trains the action expert to generate action chunks conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose, learning goal-directed action generation independently of visual cues. Stage 2 introduces a latent interface that aggregates visual and semantic representations and serves as the pretrained action expert's only visual conditioning pathway. The interface is supervised to reconstruct the terminal pose previously used to condition Stage 1, encouraging it to retain the goal-relevant spatial information needed for action generation. Across four vision-language-action and world-action architectures (Pi0.5, MolmoAct2, FAST-WAM, and ImageWAM), LIT improves overall LIBERO-Plus success by 3.87-10.70 percentage points while preserving or improving average LIBERO success. Real-world evaluations show 13.30-16.70 percentage-point gains in success aggregated across three tasks under unseen camera configurations, lighting variations, and distractors.

📄 PDF Abstract BibTeX arXiv:2609.12641

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LAMP: Latent Motion Prior-Guided Real-World Learning for Dexterous Hand Manipulation

2026-07-07 · Xinye Yang, Zhiyuan Ma, Hongze Yu, Yuanpei Chen 외 arxiv

Real-world learning for dexterous hands remains brittle because high-dimensional hand actions amplify imitation errors and make reinforcement-learning exploration prone to contact-breaking motion. While combining imitati…

Reinforcement Learning

Being-H0.7: A Latent World-Action Model from Egocentric Videos

2026-04-30 · Hao Luo, Wanpeng Zhang, Yicheng Feng, Sipeng Zheng 외 arxiv

Visual-Language-Action models (VLAs) have advanced generalist robot control by mapping multimodal observations and language instructions directly to actions, but sparse action supervision often encourages shortcut mappin…

How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?

2026-02-25 · Yingqian Cui, Zhenwei Dai, Bing He, Zhan Shi 외 arxiv

Latent reasoning has been recently proposed as a reasoning paradigm and performs multi-step reasoning through generating steps in the latent space instead of the textual space. This paradigm enables reasoning beyond disc…

Progressively Learning Heterogeneous Skills in a Unified Latent Space

2026-08-24 · Yue-Yi Zhang, Ming Gong, Linpu He, Wei-Shi Zheng 외 arxiv

We propose HetSkills, a novel framework designed to progressively learn heterogeneous skills within a unified latent space for physics-based character control. The core idea is to treat this latent space as a shared exec…

Continuous Reasoning for Vision-Language-Action

2026-05-29 · Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota arxiv

Natural language is a powerful reasoning medium for language and vision-language models, but it is mismatched to the granularity of continuous control. Text and explicit subgoals operate at task-level granularity, wherea…

Continuous Control