Cultural Binding Heads in Language Models
LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness. Using mechanistic interpretability and a factorial design on the N4 cultural appropriation benchmark from Wang et al. (2025), we identify 2-3 mid-layer attention heads per model that contribute causally to cultural binding across eight models (base and instruct versions of four architectures). Cultural binding is the process of associating a cultural item with its related identity. Knockout of the identity-to-item edges on these heads lowers the binding strength by 9-23%. The identified heads transfer from instruct to base models, suggesting that cultural binding is created during pre-training. An $α$-scaling shows a graded dose-response. Moderate amplification steering at generation ($α= 2-3$) increases cultural differentiation accuracy by 1-3 pp while leaving reasoning on culturally neutral questions mostly intact. A knowledge probing task shows that models know 3-6 times more than they act upon, indicating that the bottleneck lies in routing and not knowledge.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
I Walk the Line: Examining the Role of Gestalt Continuity in Object Binding for Vision Transformers
Object binding is a foundational process in visual cognition, during which low-level perceptual features are joined into object representations. Binding has been considered a fundamental challenge for neural networks, an…
A retrieval conditioned rebinding circuit for dynamic entity tracking in large language models
To interpret context correctly and retrieve relevant information, large language models must bind entities to their attributes and update these bindings as state changes. We analyze how LLMs implement this binding proces…
Agricultural Growth Diagnostics: Identifying the Binding Constraints and Policy Remedies for Bihar, India
Agriculture plays a significant role in economic development of the underdeveloped region. Multiple factors influence the performance of agricultural sector but a few of these have a strong bearing on its growth. We deve…
MarketingDiscovering Variable Binding Circuitry with Desiderata
Recent work has shown that computation in language models may be human-understandable, with successful efforts to localize and intervene on both single-unit features and input-output circuits. Here, we introduce an appro…
Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
Transformer-based language models display impressive reasoning-like behavior, yet remain brittle on tasks that require stable symbolic manipulation. This paper develops a unified perspective on these phenomena by interpr…