paper-with-me

홈 › Papers

MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models

2025-11-09 · Jingyu Hu, Shu Yang, Xilin Gong, Hongming Wang, Weiru Liu, Di Wang arxiv

Large Reasoning Models (LRMs) suffer from sycophantic behavior, where models tend to agree with users' incorrect beliefs and follow misinformation rather than maintain independent reasoning. This behavior undermines model reliability and poses societal risks. Mitigating LRM sycophancy requires monitoring how this sycophancy emerges during the reasoning trajectory; however, current methods mainly focus on judging based on final answers and correcting them, without understanding how sycophancy develops during reasoning processes. To address this limitation, we propose MONICA, a novel Monitor-guided Calibration framework that monitors and mitigates sycophancy during model inference at the level of reasoning steps, without requiring the model to finish generating its complete answer. MONICA integrates a sycophantic monitor that provides real-time monitoring of sycophantic drift scores during response generation with a calibrator that dynamically suppresses sycophantic behavior when scores exceed predefined thresholds. Extensive experiments across 12 datasets and 3 LRMs demonstrate that our method effectively reduces sycophantic behavior in both intermediate reasoning steps and final answers, yielding robust performance improvements.

📄 PDF Abstract BibTeX arXiv:2511.06419

Code (0)

등록된 구현이 없습니다.

Tasks

Response Generation

Similar Papers 제목 키워드 기반

Deterministic and statistical calibration of constitutive models from full-field data with parametric physics-informed neural networks

2024-05-28 · David Anton, Jendrik-Alexander Tröger, Henning Wessels, Ulrich Römer 외

The calibration of constitutive models from full-field data has recently gained increasing interest due to improvements in full-field measurement capabilities. In addition to the experimental characterization of novel ma…

Bayesian InferenceStructural Health Monitoring

HarmonicAttack: An Adaptive Cross-Domain Audio Watermark Removal

2025-11-26 · Kexin Li, Xiao Hu, Ilya Grishchenko, David Lie arxiv

The availability of high-quality, AI-generated audio raises security challenges such as misinformation campaigns and voice-cloning fraud. A key defense against the misuse of AI-generated audio is by watermarking it, so t…

Harmonica: A Self-Adaptation Exemplar for Sustainable MLOps

2026-01-17 · Ananya Halgatti, Shaunak Biswas, Hiya Bhatt, Srinivasan Rakhunathan 외 arxiv

Machine learning enabled systems (MLS) often operate in settings where they regularly encounter uncertainties arising from changes in their surrounding environment. Without structured oversight, such changes can degrade …

Time Series Regression

Correctly Modeling TX and RX Chain in (Distributed) Massive MIMO -- New Fundamental Insights on Coherency

2022-06-29 · Ronald Nissel

This letter shows that the TX and RX models commonly used in literature for downlink (distributed) massive MIMO are inaccurate, leading also to inaccurate conclusions. In particular, the Local Oscillator (LO) effect shou…

HarmonICA: Neural non-stationarity correction and source separation for motor neuron interfaces

2024-06-28 · Alexander Kenneth Clarke, Agnese Grison, Irene Mendez Guerra, Pranav Mamidanna 외

A major outstanding problem when interfacing with spinal motor neurons is how to accurately compensate for non-stationary effects in the signal during source separation routines, particularly when they cannot be estimate…