paper-with-me

홈 › Papers

Can We Trust AI Explanations? Evidence of Systematic Underreporting in Chain-of-Thought Reasoning

2025-12-25 · Deep Pankajbhai Mehta arxiv

When AI systems explain their reasoning step-by-step, practitioners often assume these explanations reveal what actually influenced the AI's answer. We tested this assumption by embedding hints into questions and measuring whether models mentioned them. In a study of over 9,000 test cases across 11 leading AI models, we found a troubling pattern: models almost never mention hints spontaneously, yet when asked directly, they admit noticing them. This suggests models see influential information but choose not to report it. Telling models they are being watched does not help. Forcing models to report hints works, but causes them to report hints even when none exist and reduces their accuracy. We also found that hints appealing to user preferences are especially dangerous-models follow them most often while reporting them least. These findings suggest that simply watching AI reasoning is not enough to catch hidden influences.

📄 PDF Abstract BibTeX arXiv:2601.00830

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations

2025-11-15 · Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan, Mingxi Yan 외 arxiv

Explanations are often promoted as tools for transparency, but they can also foster confirmation bias; users may assume reasoning is correct whenever outputs appear acceptable. We study this double-edged role of Chain-of…

Moral Scenarios

X-Blocks: Linguistic Building Blocks of Natural Language Explanations for Automated Vehicles

2026-02-02 · Ashkan Y. Zadeh, Xiaomeng Li, Andry Rakotonirainy, Ronald Schroeter 외 arxiv

Natural language explanations play a critical role in establishing trust and acceptance of automated vehicles (AVs), yet existing approaches lack systematic frameworks for analysing how humans linguistically construct dr…

Dependency Parsing

When AI Persuades: Adversarial Explanation Attacks on Human Trust in AI-Assisted Decision Making

2026-02-03 · Shutong Fan, Lan Zhang, Xiaoyong Yuan arxiv

Most adversarial threats in artificial intelligence (AI) target the computational behavior of models rather than the humans who rely on them. Yet modern AI systems increasingly operate within human decision loops, where …

Decision Making

Judgments of research co-created by generative AI: experimental evidence

2023-05-03 · Paweł Niszczota, Paul Conway

The introduction of ChatGPT has fuelled a public debate on the use of generative AI (large language models; LLMs), including its use by researchers. In the current work, we test whether delegating parts of the research p…

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

2023-05-07 · NeurIPS 2023 11 · Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman

Large Language Models (LLMs) can achieve strong performance on many tasks by producing step-by-step reasoning before giving a final output, often referred to as chain-of-thought reasoning (CoT). It is tempting to interpr…

Multiple-choice