paper-with-me

Papers

Localizing Anchoring Pathways in Language Models

2026-06-11 · Hillary N. Owusu, Sarah Wiegreffe, Naomi H. Feldman arxiv

Irrelevant numbers in a prompt can shift language model judgments, producing anchoring effects in numerical reasoning. We study where this anchor-sensitive signal is carried inside language models using a controlled multiple-choice setup with shared answer options. We define a logit-difference metric comparing the correct answer option with the answer option corresponding to the anchor, and validate that it tracks behavioral anchoring. Using attribution-based circuit localization on 7B--8B Qwen and Llama base and instruction-tuned models, we find that edge-level methods recover this signal more faithfully than node-level methods. Low- and high-anchor circuits transfer strongly within a model, suggesting shared pathway structure across anchor direction. However, sparse transfer across base and instruction-tuned variants is less reliable, indicating that post-training changes which pathways matter most. Overall, our results provide a mechanistic account of how anchoring-related decision signals are carried inside language models.

📄 PDF Abstract BibTeX arXiv:2606.12818

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

2026-08-14 · Yiderigun Borjigin, Alexander Hermann, Christian Cyron, Roland Aydin arxiv

The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that lar…

Amazon at MRP 2019: Parsing Meaning Representations with Lexical and Phrasal Anchoring

2019-11-01 · CONLL 2019 11 · Jie Cao, Yi Zhang, Adel Youssef, Vivek Srikumar

This paper describes the system submission of our team Amazon to the shared task on Cross Framework Meaning Representation Parsing (MRP) at the 2019 Conference for Computational Language Learning (CoNLL). Via extensive a…

How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations

2025-08-28 · Yoshiki Takenami, Yin Jou Huang, Yugo Murawaki, Chenhui Chu arxiv

Cognitive biases, well-studied in humans, can also be observed in LLMs, affecting their reliability in real-world applications. This paper investigates the anchoring effect in LLM-driven price negotiations. To this end, …

An Empirical Study of the Anchoring Effect in LLMs: Existence, Mechanism, and Potential Mitigations

2025-05-21 · Yiming Huang, Biquan Bie, Zuqiu Na, Weilin Ruan 외

The rise of Large Language Models (LLMs) like ChatGPT has advanced natural language processing, yet concerns about cognitive biases are growing. In this paper, we investigate the anchoring effect, a cognitive bias where …

Sequential Learning and Catastrophic Forgetting in Differentiable Resistor Networks

2026-05-02 · Maniru Ibrahim arxiv

Differentiable physical networks provide a simple setting in which learning can be studied through the interaction between trainable parameters and physical equilibrium constraints. We investigate sequential learning in …

Continual Learning