paper-with-me

홈 › Papers

Triangulation as an Acceptance Rule for Multilingual Mechanistic Interpretability

2025-12-31 · Yanan Long arxiv

Multilingual language models achieve strong aggregate performance yet often behave unpredictably across languages, scripts, and cultures. We argue that mechanistic explanations for such models should satisfy a \emph{causal} standard: claims must survive causal interventions and must \emph{cross-reference} across environments that perturb surface form while preserving meaning. We formalize \emph{reference families} as predicate-preserving variants and introduce \emph{triangulation}, an acceptance rule requiring necessity (ablating the circuit degrades the target behavior), sufficiency (patching activations transfers the behavior), and invariance (both effects remain directionally stable and of sufficient magnitude across the reference family). To supply candidate subgraphs, we adopt automatic circuit discovery and \emph{accept or reject} those candidates by triangulation. We ground triangulation in causal abstraction by casting it as an approximate transformation score over a distribution of interchange interventions, connect it to the pragmatic interpretability agenda, and present a comparative experimental protocol across multiple model families, language pairs, and tasks. Triangulation provides a falsifiable standard for mechanistic claims that filters spurious circuits passing single-environment tests but failing cross-lingual invariance.

📄 PDF Abstract BibTeX arXiv:2512.24842

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection

2026-05-21 · Yusser Al Ghussin, Daniil Gurgurov, Tanja Baeumel, Josef van Genabith 외 arxiv

Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in large language models (LLMs), but SAE-based language control remains unreliable in multilingual settings: most SAEs …

Language IdentificationMachine Translation

Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models

2025-03-08 · Thomas Winninger, Boussad ADDAD, Katarzyna Kapusta

Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targeted model, ignoring the internal mechanisms responsible for attack success or fa…

Beyond Accuracy: Introducing a Symbolic-Mechanistic Approach to Interpretable Evaluation

2026-03-06 · Reza Habibi, Darian Lee, Magy Seif El-Nasr arxiv

Accuracy-based evaluation cannot reliably distinguish genuine generalization from shortcuts like memorization, leakage, or brittle heuristics, especially in small-data regimes. In this position paper, we argue for mechan…

RomanLens: Latent Romanization and its role in Multilinguality in LLMs

2025-02-11 · Alan Saji, Jaavid Aktar Husain, Thanmay Jayakumar, Raj Dabre 외

Large Language Models (LLMs) exhibit remarkable multilingual generalization despite being predominantly trained on English-centric corpora. A fundamental question arises: how do LLMs achieve such robust multilingual capa…

Language ModelingLanguage Modelling

Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models

2025-05-22 · Ercong Nie, Helmut Schmid, Hinrich Schütze

Language confusion -- where large language models (LLMs) generate unintended languages against the user's need -- remains a critical challenge, especially for English-centric models. We present the first mechanistic inte…

BenchmarkingLanguage ModelingLanguage Modelling