paper-with-me

Papers

Multimodal Model Diffing for Feature Discovery and Control

2026-08-10 · Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr, Christian Schroeder de Witt, Constantin Venhoff, Ronald Clark hf

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.

📄 PDF Abstract BibTeX arXiv:2608.09928

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs

2026-02-12 · Thomas Jiralerspong, Trenton Bricken arxiv

Model diffing, the process of comparing models' internal representations to identify their differences, is a promising approach for uncovering safety-critical behaviors in new models. However, its application has so far …

Towards Understanding Multimodal Fine-Tuning: Spatial Features

2026-02-06 · Lachin Naghashyar, Hunar Batra, Ashkan Khakzar, Philip Torr 외 arxiv

Contemporary Vision-Language Models (VLMs) achieve strong performance on a wide range of tasks by pairing a vision encoder with a pre-trained language model, fine-tuned for visual-text inputs. Yet despite these gains, it…

Visual Grounding

List of top-venue papers related to binary diffing in the last decade

2020-04-05 · Anonymous

List of top-venue papers related to binary diffing in the last decade.

Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image Segmentation

2026-02-11 · Salma J. Ahmed, Emad A. Mohammed, Azam Asilian Bidgoli arxiv

Modern segmentation models achieve strong predictive performance but remain largely opaque, limiting our ability to diagnose failures, understand dataset shift, or intervene in a principled manner. We introduce Med-SegLe…

Medical Image Segmentation

Binary Diffing as a Network Alignment Problem via Belief Propagation

2021-12-31 · Elie Mengin, Fabrice Rossi

In this paper, we address the problem of finding a correspondence, or matching, between the functions of two programs in binary form, which is one of the most common task in binary diffing. We introduce a new formulation…