paper-with-me

Papers

Causality $\neq$ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs

2025-10-10 · Lianghuan Huang, Yingshan Chang arxiv

Mechanistic interpretability seeks to uncover how internal components of neural networks give rise to predictions. A persistent challenge, however, is disentangling two often conflated notions: decodability--the recoverability of information from hidden states--and causality--the extent to which those states functionally influence outputs. In this work, we investigate their relationship in vision transformers (ViTs) fine-tuned for object counting. Using activation patching, we test the causal role of spatial and CLS tokens by transplanting activations across clean-corrupted image pairs. In parallel, we train linear probes to assess the decodability of count information at different depths. Our results reveal systematic mismatches: middle-layer object tokens exert strong causal influence despite being weakly decodable, whereas final-layer object tokens support accurate decoding yet are functionally inert. Similarly, the CLS token becomes decodable in mid-layers but only acquires causal power in the final layers. These findings highlight that decodability and causality reflect complementary dimensions of representation--what information is present versus what is used--and that their divergence can expose hidden computational circuits.

📄 PDF Abstract BibTeX arXiv:2510.09794

Code (0)

등록된 구현이 없습니다.

Tasks

Object Counting

Similar Papers 제목 키워드 기반

Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

2026-05-01 · Gaofei Shen, Martijn Bentum, Tom Lentz, Afra Alishahi 외 arxiv

Probing is widely used to study which features can be decoded from language model representations. However, the common decoding probe approach has two limitations that we aim to solve with our new encoding probe approach…

Toward Machine Interpreting: Lessons from Human Interpreting Studies

2025-08-11 · Matthias Sperber, Maureen de Seyssel, Jiajun Bao, Matthias Paulik arxiv

Current speech translation systems, while having achieved impressive accuracies, are rather static in their behavior and do not adapt to real-world situations in ways human interpreters do. In order to improve their prac…

Machine Translation

The MeVer DeepFake Detection Service: Lessons Learnt from Developing and Deploying in the Wild

2022-04-27 · Spyridon Baxevanakis, Giorgos Kordopatis-Zilos, Panagiotis Galopoulos, Lazaros Apostolidis 외

Enabled by recent improvements in generation methodologies, DeepFakes have become mainstream due to their increasingly better visual quality, the increase in easy-to-use generation tools and the rapid dissemination throu…

DeepFake DetectionFace Swapping

From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability

2026-08-09 · Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou arxiv

How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric…

Linear and nonlinear causality in financial markets

2023-12-18 · Haochun Ma, Davide Prosperino, Alexander Haluszczynski, Christoph Räth

Identifying and quantifying co-dependence between financial instruments is a key challenge for researchers and practitioners in the financial industry. Linear measures such as the Pearson correlation are still widely use…

Causal InferenceManagementPAIR TRADING