paper-with-me

홈 › Papers

Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries

2025-02-27 · Tianyi Lorena Yan, Robin Jia

To answer one-to-many factual queries (e.g., listing cities of a country), a language model (LM) must simultaneously recall knowledge and avoid repeating previous answers. How are these two subtasks implemented and integrated internally? Across multiple datasets and models, we identify a promote-then-suppress mechanism: the model first recalls all answers, and then suppresses previously generated ones. Specifically, LMs use both the subject and previous answer tokens to perform knowledge recall, with attention propagating subject information and MLPs promoting the answers. Then, attention attends to and suppresses previous answer tokens, while MLPs amplify the suppression signal. Our mechanism is corroborated by extensive experimental evidence: in addition to using early decoding and causal tracing, we analyze how components use different tokens by introducing both Token Lens, which decodes aggregated attention updates from specified tokens, and a knockout method that analyzes changes in MLP outputs after removing attention to specified tokens. Overall, we provide new insights into how LMs' internal components interact with different input tokens to support complex factual recall. Code is available at https://github.com/Lorenayannnnn/how-lms-answer-one-to-many-factual-queries.

📄 PDF Abstract BibTeX arXiv:2502.20475

Code (1)

lorenayannnnn/how-lms-answer-one-to-many-factual-queries 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Measuring Safety Alignment Effects in Autonomous Security Agents

2026-05-19 · Isaac David, Arthur Gervais arxiv

Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Single-turn refusal benchmarks cannot answer this question: security agen…

Multi-label Iterated Learning for Image Classification with Label Ambiguity

2021-11-23 · CVPR 2022 1 · Sai Rajeswar, Pau Rodriguez, Soumye Singhal, David Vazquez 외

Transfer learning from large-scale pre-trained models has become essential for many computer vision tasks. Recent studies have shown that datasets like ImageNet are weakly labeled since images with multiple object classe…

Classificationimage-classificationImage ClassificationMulti-Label Learning+1

Reasoning about Intent for Ambiguous Requests

2025-11-13 · Irina Saparina, Mirella Lapata arxiv

Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating safety risks when that interpretation is wrong. We propose generating a single stru…

Conversational Question AnsweringReinforcement LearningSemantic Parsing

How Language Models Process Negation

2026-05-04 · Zhejian Zhou, Tianyi Zhou, Robin Jia, Jonathan May arxiv

We study how Large Language Models (LLMs) process negation mechanistically. First, we establish that even though open-weight models often provide wrong answers to questions involving negation, they do possess internal co…

Overcoming Label Ambiguity with Multi-label Iterated Learning

2021-09-29 · Sai Rajeswar Mudumba, Pau Rodriguez, Soumye Singhal, David Vazquez 외

Transfer learning from ImageNet pre-trained models has become essential for many computer vision tasks. Recent studies have shown that ImageNet includes label ambiguity, where images with multiple object classes present …

Multi-Label LearningTransfer Learning