paper-with-me

Papers

Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

2026-05-01 · Gaofei Shen, Martijn Bentum, Tom Lentz, Afra Alishahi, Grzegorz Chrupała arxiv

Probing is widely used to study which features can be decoded from language model representations. However, the common decoding probe approach has two limitations that we aim to solve with our new encoding probe approach: contributions of different features to model representations cannot be directly compared, and feature correlations can affect probing results. We present an Encoding Probe that reverses this direction and reconstructs internal representations of models using interpretable features. We evaluate this method on text and speech transformer models, using feature sets spanning acoustics, phonetics, syntax, lexicon, and speaker identity. Our results suggest that speaker-related effects vary strongly across different training objectives and datasets, while syntactic and lexical features contribute independently to reconstruction. These results show that the Encoding Probe provides a complementary perspective on interpreting model representations beyond decodability.

📄 PDF Abstract BibTeX arXiv:2605.00607

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability

2026-08-09 · Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou arxiv

How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric…

Dissociating Decodability and Causal Use in Bracket-Sequence Transformers

2026-04-24 · Aryan Sharma, Cutter Dawes, Shivam Raval arxiv

When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residual stream, and in stack-like attention p…

Probing in the Wild: A Case Study of Self-Supervised Speech Representations on Mandarin Sub-dialects with Unsupervised Articulatory Analysis

2026-06-24 · Shu Shang, Fuliang Weng, Zeqian Hu, Yaqian Zhou arxiv

While self-supervised speech models have achieved strong performance across speech tasks, relatively little is known about how their internal phonetic representations behave under fine-grained dialect variation. Existing…

CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

2026-08-21 · Bokai Zhao, Yiyang Zhang, Hanqing Chao, Yawei Ma 외 arxiv

Pathology foundation models (PFMs) are increasingly used as general-purpose backbones, yet existing benchmarks cannot systematically diagnose their whole-slide cellular representation capabilities, including the decodabi…

Domain Generalization

Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

2026-08-18 · Man Liang, Xinzhao Cheng, Faizan Wajid arxiv

Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric r…