paper-with-me

홈 › Papers

Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs

2026-06-14 · Nafiseh Nikeghbal, Amir Hossein Kargaran, Shaghayegh Kolli, Jana Diesner arxiv

Standard accuracy benchmarks are designed to test how closely large language models (LLMs) approach correct answers, but are not suitable for testing whether LLMs stick with a correct answer when that answer is challenged by a plausible counter-argument. We introduce a controlled protocol for evaluating answer stability: after a model answers a multiple-choice question correctly, we challenge the model's answer with a coherent argument for an incorrect option and measure whether the model flips. The setup a) isolates argumentative content from overt social pressure and b) varies argument length, self-attribution, and cross-model source. Across seven frontier models and 57 MMLU subjects, flip rates range from 17.5% to 97.3%, revealing large differences in stability that are not captured by accuracy metrics alone. We find that self-attribution consistently increases flip rates (mean +7.1pp, up to +18.7pp). Also, pooling wrong-answer arguments across models and selecting the most effective one per question yields stronger adversarial challenges than relying on any single source model. We further construct MaxFlip, a curated challenge set that amplifies flips by up to +23.6pp over standard self-generated challenges. We release the protocol, challenge records, and MaxFlip to support stability evaluation alongside standard accuracy benchmarks. Materials are available at https://github.com/nafisenik/WhoFlips and https://hf.co/datasets/nafisehNik/WhoFlips.

📄 PDF Abstract BibTeX arXiv:2606.16011

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate

2026-05-30 · Xiqi Hao, Zengqing Wu, Yu-Xuan Qiu, Chuan Xiao 외 arxiv

Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that convergence reflects genuine deliberation or social compliance. We sho…

Counterargument for Critical Thinking as Judged by AI and Humans

2026-05-06 · Tosin Adewumi, Marcus Liwicki, Foteini Simistira Liwicki, Lama Alkhaled 외 arxiv

This intervention study investigates the use of counterarguments in writing for critical thinking by students in the context of Generative AI (GenAI). This is especially as risks of cheating and cognitive offloading exis…

Benchmarking Multi-turn Medical Diagnosis: Hold, Lure, and Self-Correction

2026-04-06 · Jinrui Fang, Runhan Chen, Xu Yang, Jian Yu 외 arxiv

Large language models (LLMs) achieve high accuracy in medical diagnosis when all clinical information is provided in a single turn, yet how they behave under multi-turn evidence accumulation closer to real clinical reaso…

Medical Diagnosis

Argument Harvesting Using Chatbots

2018-05-11 · Lisa A. Chalaguine, Anthony Hunter, Henry W. W. Potts, Fiona L. Hamilton

Much research in computational argumentation assumes that arguments and counterarguments can be obtained in some way. Yet, to improve and apply models of argument, we need methods for acquiring them. Current approaches i…

Argument MiningChatbot

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models

2026-06-24 · Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli arxiv

Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by eme…

Visual Reasoning