paper-with-me

Papers

Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions

2025-05-21 · Sasha Boguraev, Christopher Potts, Kyle Mahowald

Large Language Models (LLMs) have emerged as powerful sources of evidence for linguists seeking to develop theories of syntax. In this paper, we argue that causal interpretability methods, applied to LLMs, can greatly enhance the value of such evidence by helping us characterize the abstract mechanisms that LLMs learn to use. Our empirical focus is a set of English filler-gap dependency constructions (e.g., questions, relative clauses). Linguistic theories largely agree that these constructions share many properties. Using experiments based in Distributed Interchange Interventions, we show that LLMs converge on similar abstract analyses of these constructions. These analyses also reveal previously overlooked factors -- relating to frequency, filler type, and surrounding context -- that could motivate changes to standard linguistic theory. Overall, these results suggest that mechanistic, internal analyses of LLMs can push linguistic theory forward.

📄 PDF Abstract BibTeX arXiv:2505.16002

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

MetaCaDI: A Meta-Learning Framework for Causal Discovery from Multiple Environments with Unknown Interventions

2025-10-25 · Hans Jarett Ong, Yoichi Chikahara, Tomoharu Iwata arxiv

Uncovering the causal mechanisms of complex real-world systems remains a significant challenge, as these systems often entail high data collection costs and involve unknown interventions. We introduce MetaCaDI, the first…

Bilevel Optimization

Transferring Information Across Interventions in Causal Bayesian Optimization

2026-05-31 · Mohammad Ali Javidian arxiv

Bayesian optimization is a popular way to optimize expensive systems, where every experiment, simulation, or intervention costs time or money. In its standard form, it treats the variables we control as plain inputs to a…

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models

2025-09-18 · Qidong Wang, Junjie Hu, Ming Jiang arxiv

Recent advances in causal interpretability have extended from language models to vision-language models (VLMs), seeking to reveal their internal mechanisms through input interventions. While textual interventions often t…

Emergent Causal-Geometric Dynamics Across Depth in Large Language Models

2026-02-04 · Shahar Haim, Daniel C McNamee arxiv

Geometric analyses of large language model (LLM) representations reveal structured variation across depth but remain fundamentally correlational with respect to token prediction formation. Meanwhile, causal interventions…

Scalable Contrastive Causal Discovery under Unknown Soft Interventions

2026-03-03 · Mingxuan Zhang, Khushi Desai, Sopho Kevlishvili, Elham Azizi arxiv

Observational causal discovery is only identifiable up to the Markov equivalence class. While interventions can reduce this ambiguity, in practice interventions are often soft with multiple unknown targets. In many reali…