paper-with-me

홈 › Papers

Tracing Facts or just Copies? A critical investigation of the Competitions of Mechanisms in Large Language Models

2025-07-16 · Dante Campregher, Yanxu Chen, Sander Hoffman, Maria Heuss arxiv

This paper presents a reproducibility study examining how Large Language Models (LLMs) manage competing factual and counterfactual information, focusing on the role of attention heads in this process. We attempt to reproduce and reconcile findings from three recent studies by Ortu et al., Yu, Merullo, and Pavlick and McDougall et al. that investigate the competition between model-learned facts and contradictory context information through Mechanistic Interpretability tools. Our study specifically examines the relationship between attention head strength and factual output ratios, evaluates competing hypotheses about attention heads' suppression mechanisms, and investigates the domain specificity of these attention patterns. Our findings suggest that attention heads promoting factual output do so via general copy suppression rather than selective counterfactual suppression, as strengthening them can also inhibit correct facts. Additionally, we show that attention head behavior is domain-dependent, with larger models exhibiting more specialized and category-sensitive patterns.

📄 PDF Abstract BibTeX arXiv:2507.11809

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Tracing the Origin of Adversarial Attack for Forensic Investigation and Deterrence

2022-12-31 · ICCV 2023 1 · Han Fang, Jiyi Zhang, Yupeng Qiu, Ke Xu 외

Deep neural networks are vulnerable to adversarial attacks. In this paper, we take the role of investigators who want to trace the attack and identify the source, that is, the particular model which the adversarial examp…

Adversarial Attack

Tracing the Roots of Facts in Multilingual Language Models: Independent, Shared, and Transferred Knowledge

2024-03-08 · Xin Zhao, Naoki Yoshinaga, Daisuke Oba

Acquiring factual knowledge for language models (LMs) in low-resource languages poses a serious challenge, thus resorting to cross-lingual transfer in multilingual LMs (ML-LMs). In this study, we ask how ML-LMs acquire a…

Cross-Lingual TransferKnowledge ProbingRepresentation Learning

Robust Identity Perceptual Watermark Against Deepfake Face Swapping

2023-11-02 · Tianyi Wang, Mengxiao Huang, Harry Cheng, Bin Ma 외

Notwithstanding offering convenience and entertainment to society, Deepfake face swapping has caused critical privacy issues with the rapid development of deep generative models. Due to imperceptible artifacts in high-qu…

DecoderFace Swapping

SEMDR: A Semantic-Aware Dual Encoder Model for Legal Judgment Prediction with Legal Clue Tracing

2024-08-19 · Pengjie Liu, Wang Zhang, Yulong Ding, Xuefeng Zhang 외

Legal Judgment Prediction (LJP) aims to form legal judgments based on the criminal fact description. However, researchers struggle to classify confusing criminal cases, such as robbery and theft, which requires LJP model…

Representation LearningSentence

On the Interpretability of Deep Learning Based Models for Knowledge Tracing

2021-01-27 · Xinyi Ding, Eric C. Larson

Knowledge tracing allows Intelligent Tutoring Systems to infer which topics or skills a student has mastered, thus adjusting curriculum accordingly. Deep Learning based models like Deep Knowledge Tracing (DKT) and Dynami…

Decision MakingKnowledge Tracing