paper-with-me

Papers

Feedback Indices to Evaluate LLM Responses to Rebuttals for Multiple Choice Type Questions

2026-01-02 · Justin C. Dunlap, Anne-Simone Parent, Ralf Widenhorn arxiv

We present a systematic framework of indices designed to characterize Large Language Model (LLM) responses when challenged with rebuttals during a chat. Assessing how LLMs respond to user dissent is crucial for understanding their reliability and behavior patterns, yet the complexity of human-LLM interactions makes systematic evaluation challenging. Our approach employs a fictitious-response rebuttal method that quantifies LLM behavior when presented with multiple-choice questions followed by deliberate challenges to their fictitious previous response. The indices are specifically designed to detect and measure what could be characterized as sycophantic behavior (excessive agreement with user challenges) or stubborn responses (rigid adherence to the fictitious response in the chat history) from LLMs. These metrics allow investigation of the relationships between sycophancy, stubbornness, and the model's actual mastery of the subject matter. We demonstrate the utility of these indices using two physics problems as test scenarios with various OpenAI models. The framework is intentionally generalizable to any multiple-choice format question, including on topics without universally accepted correct answers. Our results reveal measurable differences across OpenAI model generations, with trends indicating that newer models and those employing greater "Reasoning Effort" exhibit reduced sycophantic behavior. The FR pairing method combined with our proposed indices provides a practical, adaptable toolkit for systematically comparing LLM dialogue behaviors across different models and contexts.

📄 PDF Abstract BibTeX arXiv:2601.03285

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025

2025-04-13 · Nitya Thakkar, Mert Yuksekgonul, Jake Silberg, Animesh Garg 외

Peer review at AI conferences is stressed by rapidly rising submission volumes, leading to deteriorating review quality and increased author dissatisfaction. To address these issues, we developed Review Feedback Agent, a…

SycEval: Evaluating LLM Sycophancy

2025-02-12 · Aaron Fanous, Jacob Goldberg, Ank A. Agarwal, Joanna Lin 외

Large language models (LLMs) are increasingly applied in educational, clinical, and professional settings, but their tendency for sycophancy -- prioritizing user agreement over independent reasoning -- poses risks to rel…

Model Optimization

ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation

2026-09-08 · Yiling Ma, Yilun Zhao, Sihong Wu, Ziyu Chen 외 hf

As LLMs are increasingly used for pre-submission self-review, there is growing demand for feedback that not only identifies weaknesses but also guides authors toward concrete revisions. We study this as Actionable Peer-r…

A Dataset of General-Purpose Rebuttal

2019-09-01 · IJCNLP 2019 11 · Matan Orbach, Yonatan Bilu, Ariel Gera, Yoav Kantor 외

In Natural Language Understanding, the task of response generation is usually focused on responses to short texts, such as tweets or a turn in a dialog. Here we present a novel task of producing a critical response to a …

Natural Language UnderstandingResponse Generation

FeedbackMap: a tool for making sense of open-ended survey responses

2023-06-26 · Doug Beeferman, Nabeel Gillani

Analyzing open-ended survey responses is a crucial yet challenging task for social scientists, non-profit organizations, and educational institutions, as they often face the trade-off between obtaining rich data and the …

Survey