paper-with-me

홈 › Papers

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control

2026-02-13 · William Chen, Jagdeep Singh Bhatia, Catherine Glossop, Nikhil Mathihalli, Ria Doshi, Andy Tang, Danny Driess, Karl Pertsch, Sergey Levine arxiv

Pretrained vision-language models (VLMs) can make semantic and visual inferences across diverse settings, providing valuable common-sense priors for robotic control. However, effectively grounding this knowledge in robot behaviors remains an open challenge. Prior methods often employ a hierarchical approach where VLMs reason over high-level commands to be executed by separate low-level policies, e.g., vision-language-action models (VLAs). The interface between VLMs and VLAs is usually natural language task instructions, which fundamentally limits how much VLM reasoning can steer low-level behavior. We thus introduce Steerable Policies: VLAs trained on rich synthetic commands at various levels of abstraction, like subtasks, motions, and grounded pixel coordinates. By improving low-level controllability, Steerable Policies can unlock pretrained knowledge in VLMs, enabling improved task generalization. We demonstrate this benefit by controlling our Steerable Policies with both a learned high-level embodied reasoner and an off-the-shelf VLM prompted to reason over command abstractions via in-context learning. Across extensive real-world manipulation experiments, these two novel methods outperform prior embodied reasoning VLAs and VLM-based hierarchical baselines, including on challenging generalization and long-horizon tasks. Website: steerable-policies.github.io

📄 PDF Abstract BibTeX arXiv:2602.13193

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

2026-05-13 · Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang 외 arxiv

In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to guide robot policies with explicit visu…

Mechanistic interpretability for steering vision-language-action models

2025-08-30 · Bear Häon, Kaylene Stocking, Ian Chuang, Claire Tomlin arxiv

Vision-Language-Action (VLA) models are a promising path to realizing generalist embodied agents that can quickly adapt to new tasks, modalities, and environments. However, methods for interpreting and steering VLAs fall…

VAMOS: A Hierarchical Vision-Language-Action Model for Capability-Modulated and Steerable Navigation

2025-10-23 · Mateo Guaman Castro, Sidharth Rajagopal, Daniel Gorbatov, Matt Schmittle 외 arxiv

A fundamental challenge in robot navigation lies in learning policies that generalize across diverse environments while conforming to the unique physical constraints and capabilities of a specific embodiment (e.g., quadr…

Robot Navigation

DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning

2026-01-22 · Junha Lee, Eunha Park, Minsu Cho arxiv

Language-driven dexterous grasp generation requires the models to understand task semantics, 3D geometry, and complex hand-object interactions. While vision-language models have been applied to this problem, existing app…

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

2026-05-26 · Xintong Hu, Xuhong Huang, Jinyu Zhang, Yutong Yao 외 arxiv

Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tasks should be executed. However, existing robot datasets usually pair t…