paper-with-me

Papers

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

2026-07-16 · Jihoon Hong, Julian Skifstad, Qiyue Dai, Alice Chan, Glen Chou arxiv

World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This motivates the use of contrastive activation directions for training-free WAM steering. We also show that local linearity in WAM activation dynamics enables efficient feedback steering via model-based optimal control, yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller. Via mechanistic evaluations, we predict strong steerability in the Cosmos-Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations over unsteered and prompt steering baselines.

📄 PDF Abstract BibTeX arXiv:2607.14943

Code (1)

umerjavaidkh/AI_Research_Collection

Similar Papers 제목 키워드 기반

Mechanistic interpretability for steering vision-language-action models

2025-08-30 · Bear Häon, Kaylene Stocking, Ian Chuang, Claire Tomlin arxiv

Vision-Language-Action (VLA) models are a promising path to realizing generalist embodied agents that can quickly adapt to new tasks, modalities, and environments. However, methods for interpreting and steering VLAs fall…

Interpretability without actionability: mechanistic methods cannot correct language model errors despite near-perfect internal representations

2026-03-18 · Sanjay Basu, Sadiq Y. Patel, Parth Sheth, Bhairavi Muralidharan 외 arxiv

Language models encode task-relevant knowledge in internal representations that far exceeds their output performance, but whether mechanistic interpretability methods can bridge this knowledge-action gap has not been sys…

Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations

2025-11-27 · Chancharik Mitra, Yusen Luo, Raj Saravanan, Dantong Niu 외 arxiv

Vision-Language Action (VLAs) models promise to extend the remarkable success of vision-language models (VLMs) to robotics. Yet, unlike VLMs in the vision-language domain, VLAs for robotics require finetuning to contend …

When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability

2026-01-06 · Raphael Ronge, Markus Maier, Frederick Eberhardt arxiv

Recent work by Anthropic on Mechanistic interpretability claims to understand and control Large Language Models by extracting human-interpretable features from their neural activation patterns using sparse autoencoders (…

To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models

2025-10-15 · Anna Hedström, Salim I. Amoukou, Tom Bewley, Saumitra Mishra 외 arxiv

We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely o…