paper-with-me

Papers

Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models

2026-04-02 · Ayush Rajesh Jhaveri, Anthony GX-Chen, Ilia Sucholutsky, Eunsol Choi arxiv

Confirmation bias, the tendency to seek evidence that supports rather than challenges one's belief, hinders one's reasoning ability. We examine whether large language models (LLMs) exhibit confirmation bias by adapting the rule-discovery study from human psychology: given a sequence of three numbers (a "triple"), an agent engages in an interactive feedback loop where it (1) proposes a new triple, (2) receives feedback on whether it satisfies the hidden rule, and (3) guesses the rule. Across eleven LLMs of multiple families and scales, we find that LLMs exhibit confirmation bias, often proposing triples to confirm their hypothesis rather than trying to falsify it. This leads to slower and less frequent discovery of the hidden rule. We further explore intervention strategies (e.g., encouraging the agent to consider counter examples) developed for humans. We find prompting LLMs with such instruction consistently decreases confirmation bias in LLMs, improving rule discovery rates from 42% to 56% on average. Lastly, we mitigate confirmation bias by distilling intervention-induced behavior into LLMs, showing promising generalization to a new task, the Blicket test. Our work shows that confirmation bias is a limitation of LLMs in hypothesis exploration, and that it can be mitigated via injecting interventions designed for humans.

📄 PDF Abstract BibTeX arXiv:2604.02485

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games

2026-06-03 · Leonardo Bertolazzi, Katya Tentori, Raffaella Bernardi arxiv

Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks. Yet whether these systems can effectively engage in forms of inductive reasoning relevant to scientific discovery remains a…

DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark

2026-08-31 · Jayanta Sadhu, Sayem Shahad, Kenneth Marino arxiv

Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of newer evidence. Although recent studies have examined language-model b…

SemiReward: A General Reward Model for Semi-supervised Learning

2023-10-04 · Siyuan Li, Weiyang Jin, Zedong Wang, Fang Wu 외

Semi-supervised learning (SSL) has witnessed great progress with various improvements in the self-training framework with pseudo labeling. The main challenge is how to distinguish high-quality pseudo labels against the c…

Few-Shot Image ClassificationImage ClassificationPseudo LabelSemi-supervised Audio Classification+3

MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation

2025-09-26 · Xinping Lei, Tong Zhou, Yubo Chen, Kang Liu 외 arxiv

Large Language Models (LLMs) hold substantial potential for accelerating academic ideation but face critical challenges in grounding ideas and mitigating confirmation bias for further refinement. We propose integrating m…

Knowledge Graphs

Confirmation Bias in Generative AI Chatbots: Mechanisms, Risks, Mitigation Strategies, and Future Research Directions

2025-04-12 · Yiran Du

This article explores the phenomenon of confirmation bias in generative AI chatbots, a relatively underexamined aspect of AI-human interaction. Drawing on cognitive psychology and computational linguistics, it examines h…

Chatbot