paper-with-me

홈 › Papers

Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units

2026-01-29 · Jianhui Chen, Yuzhang Luo, Liangming Pan arxiv

While Mechanistic Interpretability has identified interpretable circuits in LLMs, their causal origins in training data remain elusive. We introduce Mechanistic Data Attribution (MDA), a scalable framework that employs Influence Functions to trace interpretable units back to specific training samples. Through extensive experiments on the Pythia family, we causally validate that targeted intervention--removing or augmenting a small fraction of high-influence samples--significantly modulates the emergence of interpretable heads, whereas random interventions show no effect. Our analysis reveals that repetitive structural data (e.g., LaTeX, XML) acts as a mechanistic catalyst. Furthermore, we observe that interventions targeting induction head formation induce a concurrent change in the model's in-context learning (ICL) capability. This provides direct causal evidence for the long-standing hypothesis regarding the functional link between induction heads and ICL. Finally, we propose a mechanistic data augmentation pipeline that consistently accelerates circuit convergence across model scales, providing a principled methodology for steering the developmental trajectories of LLMs.

📄 PDF Abstract BibTeX arXiv:2601.21996

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies

2026-06-28 · Reza Habibi, Darian Lee, Magy Seif El-Nasr arxiv

While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how training data shapes the high-level behavioral decisions a model learns to make. T…

Deepfake Forensic Analysis: Source Dataset Attribution and Legal Implications of Synthetic Media Manipulation

2025-05-16 · Massimiliano Cassia, Luca Guarnera, Mirko Casu, Ignazio Zangara 외

Synthetic media generated by Generative Adversarial Networks (GANs) pose significant challenges in verifying authenticity and tracing dataset origins, raising critical concerns in copyright enforcement, privacy protectio…

Binary ClassificationFace Swapping

Origin Tracing and Detecting of LLMs

2023-04-27 · Linyang Li, Pengyu Wang, Ke Ren, Tianxiang Sun 외

The extraordinary performance of large language models (LLMs) heightens the importance of detecting whether the context is generated by an AI system. More importantly, while more and more companies and institutions relea…

Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation

2025-09-17 · Baolei Zhang, Haoran Xin, Yuxi Chen, Zhuqing Liu 외 arxiv

Retrieval-Augmented Generation (RAG) integrates external knowledge into large language models to improve response quality. However, recent work has shown that RAG systems are highly vulnerable to poisoning attacks, where…

The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing

2026-03-21 · Yuan Cao, Mingyang Wang, Hinrich Schütze arxiv

Large language models (LLMs) are increasingly used as knowledge bases, but keeping them up to date requires targeted knowledge editing (KE). However, it remains unclear how edits are implemented inside the model once app…

knowledge editing