paper-with-me

홈 › Papers

Interpreting Brain Responses to Language with Sparse Features from Language Models

2026-06-05 · Michael A. Lepori, Kendrick Kay, Greta Tuckute arxiv

A central goal of cognitive neuroscience is to characterize the features that are represented by human language cortex. Artificial language models (LMs) have emerged as a powerful tool to address this challenge, but studies relating biological and artificial representations are often criticized as relating one black box to another. The present work introduces Augmented Sparse Encoding Models, an encoding framework that replaces dense LM hidden states with hierarchically-organized sparse autoencoder (SAE) features, while explicitly including surprisal as a predictor. Using this approach, we (i) produce interpretations of neural responses and (ii) test whether model-brain alignment reflects primary or idiosyncratic variation in LM representations. Using a high-field 7T fMRI dataset of eight participants listening to 200 linguistically diverse sentences, we first validate our modeling framework by recovering previous interpretations of voxel populations tuned to processing difficulty and meaning abstractness. We then interpret a previously-uncharacterized (but reliable) voxel population and find that it is tuned to people-related content. Next, we show that the fronto-temporal human language network is predicted by a common set of features across its constituent regions, but find that frontal regions are relatively well-explained by surprisal alone, even in the absence of LM-based features. Finally, we show that brain responses during language processing are not merely predictable from an arbitrary set of LM features. Rather, brain responses are best explained by the features that tend to capture the most general information encoded in LM representations, suggesting a nontrivial correspondence between brain and LM language representation.

📄 PDF Abstract BibTeX arXiv:2606.06857

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SAFER: Probing Safety in Reward Models with Sparse Autoencoder

2025-07-01 · Wei Shi, Ziyuan Xie, Sihang Li, Xiang Wang arxiv

Reinforcement learning from human feedback (RLHF) is a key paradigm for aligning large language models (LLMs) with human values, yet the reward models at its core remain largely opaque. In this work, we present Sparse Au…

Reinforcement Learning

Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography

2026-05-21 · Dongxin Guo, Jikun Wu, Siu Ming Yiu arxiv

Intermediate layers of large language models (LLMs) best predict human brain responses to language, one of the most robust findings in computational neurolinguistics, yet why remains mechanistically unexplained. We addre…

LinBridge: A Learnable Framework for Interpreting Nonlinear Neural Encoding Models

2024-10-26 · Xiaohui Gao, Yue Cheng, Peiyang Li, Yijie Niu 외

Neural encoding of artificial neural networks (ANNs) links their computational representations to brain responses, offering insights into how the brain processes information. Current studies mostly use linear encoding mo…

Self-Supervised Learning

Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders

2025-02-21 · Xuansheng Wu, Jiayi Yuan, Wenlin Yao, Xiaoming Zhai 외

Large language models (LLMs) excel at handling human queries, but they can occasionally generate flawed or unexpected responses. Understanding their internal states is crucial for understanding their successes, diagnosin…

Brain Dissection: fMRI-trained Networks Reveal Spatial Selectivity in the Processing of Natural Images

2023-09-21 · NeurIPS 2023 11

The alignment between deep neural network (DNN) features and cortical responses currently provides the most accurate quantitative explanation for higher visual areas. At the same time, these model features have been crit…