paper-with-me

Papers

Event-Grounded Sparse Autoencoders for Vision-Language-Action Policies

2026-05-17 · Xinchen Jin, Aditya Chatterjee, Pranav Kumar, Rohan Paleja arxiv

Vision-Language-Action (VLA) policies translate language and visual inputs into robot actions, where their hidden representations directly shape closed-loop behavior. However, mechanistic interpretability tools from language and vision-language models do not transfer cleanly to VLAs: outputs are robot actions rather than human-readable tokens, and interventions can only be tested via expensive closed-loop rollouts. We propose an event-grounded interpretability pipeline that anchors SAE feature analysis to behavioral events rather than text contexts. End-effector keyframes are clustered within each task using visual, state, and temporal cues, linking SAE features to behaviorally salient events and, via optional VLM annotations, to semantic context. To our knowledge, our pipeline is among the first to ground SAE-based VLA analysis in closed-loop behavioral events. Across two simulation architectures and a real-robot study, event-grounded ranking yields the strongest causal effects on OpenVLA and transfers to the continuous action chunks of $π_{0.5}$. SAE is a sparse but imperfect intervention basis: usability varies with architecture and intervention site, and aggressive intervention reveals safety and interpretability limits. Overall, event-grounded SAE analysis emerges as a practical starting point for behavior-anchored VLA interpretability, motivating future work on SAE features beyond action-aligned coordinates, finer-grained closed-loop evaluation, and safe interventions for high-stakes VLA deployments. Code is available at \url{https://github.com/xc-j/Event-SAE}.

📄 PDF Abstract BibTeX arXiv:2605.17204

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sparse Coding with Multi-Layer Decoders using Variance Regularization

2021-12-16 · Katrina Evtimova, Yann Lecun

Sparse representations of images are useful in many computer vision applications. Sparse coding with an $l_1$ penalty and a learned linear dictionary requires regularization of the dictionary to prevent a collapse in the…

DecoderDenoising

Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

2026-06-23 · Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos 외 arxiv

Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather th…

Semantic correspondence

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models

2026-05-21 · Dimitrios Damianos, Leon Voukoutis, Georgios Skyrianos, Vassilis Katsouros 외 arxiv

Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood. Existing interpretability work on VLMs uses Sparse Autoencoders (SAE…

Multimodal ReasoningVisual Grounding

Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models

2025-04-03 · Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge Belongie 외

Sparse Autoencoders (SAEs) have recently been shown to enhance interpretability and steerability in Large Language Models (LLMs). In this work, we extend the application of SAEs to Vision-Language Models (VLMs), such as …

Evaluation of Vision-LLMs in Surveillance Video

2025-10-27 · Pascal Benschop, Cristian Meo, Justin Dauwels, Jelte P. Mense arxiv

The widespread use of cameras in our society has created an overwhelming amount of video data, far exceeding the capacity for human monitoring. This presents a critical challenge for public safety and security, as the ti…

Action RecognitionSpatial Reasoning