paper-with-me

홈 › Papers

Characterizing Mechanisms for Factual Recall in Language Models

2023-10-24 · Qinan Yu, Jack Merullo, Ellie Pavlick

Language Models (LMs) often must integrate facts they memorized in pretraining with new information that appears in a given context. These two sources can disagree, causing competition within the model, and it is unclear how an LM will resolve the conflict. On a dataset that queries for knowledge of world capitals, we investigate both distributional and mechanistic determinants of LM behavior in such situations. Specifically, we measure the proportion of the time an LM will use a counterfactual prefix (e.g., "The capital of Poland is London") to overwrite what it learned in pretraining ("Warsaw"). On Pythia and GPT2, the training frequency of both the query country ("Poland") and the in-context city ("London") highly affect the models' likelihood of using the counterfactual. We then use head attribution to identify individual attention heads that either promote the memorized answer or the in-context answer in the logits. By scaling up or down the value vector of these heads, we can control the likelihood of using the in-context answer on new data. This method can increase the rate of generating the in-context answer to 88\% of the time simply by scaling a single head at runtime. Our work contributes to a body of evidence showing that we can often localize model behaviors to specific components and provides a proof of concept for how future methods might control model behavior dynamically at runtime.

📄 PDF Abstract BibTeX arXiv:2310.15910

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactual

Methods 이 논문이 사용한 방법론

Pythia Pythia is a suite of decoder-only autoregressive language models all trained on public data seen in the exact same order and ranging in size from 70M to 12B parameters. The…

Similar Papers 제목 키워드 기반

How Do Multilingual Models Remember? Investigating Multilingual Factual Recall Mechanisms

2024-10-18 · Constanza Fierro, Negar Foroutan, Desmond Elliott, Anders Søgaard

Large Language Models (LLMs) store and retrieve vast amounts of factual knowledge acquired during pre-training. Prior research has localized and identified mechanisms behind knowledge recall; however, it has primarily fo…

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?

2026-05-21 · Luca Modica, Filip Landin, Mehrdad Farahani, Livia Qian 외 arxiv

In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges about how model-internal mechanisms are similar and different when oper…

Text to Speech

Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall Mechanisms

2025-09-10 · Minyeong Choe, Haehyun Cho, Changho Seo, Hyunil Kim arxiv

Understanding how Transformer-based language models store and retrieve factual associations is critical for improving interpretability and enabling targeted model editing. Prior work, primarily on GPT-style models, has i…

Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval

2025-12-02 · Constantin Venhoff, Ashkan Khakzar, Sonia Joseph, Philip Torr 외 arxiv

Training vision language models (VLMs) aims to align visual representations from a vision encoder with the textual representations of a pretrained large language model (LLM). However, many VLMs exhibit reduced factual re…

Entity Resolution

A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia

2023-12-04 · Giovanni Monea, Maxime Peyrard, Martin Josifoski, Vishrav Chaudhary 외

Large language models (LLMs) have an impressive ability to draw on novel information supplied in their context. Yet the mechanisms underlying this contextual grounding remain unknown, especially in situations where conte…

counterfactualLanguage ModelingLanguage ModellingRetrieval+1