paper-with-me

홈 › Papers

Evaluating Evidence Grounding Under User Pressure in Instruction-Tuned Language Models

2026-03-20 · Sai Koneru, Elphin Joe, Christine Kirchhoff, Jian Wu, Sarah Rajtmajer arxiv

In contested domains, instruction-tuned language models must balance user-alignment pressures against faithfulness to the in-context evidence. To evaluate this tension, we introduce a controlled epistemic-conflict framework grounded in the U.S. National Climate Assessment. We conduct fine-grained ablations over evidence composition and uncertainty cues across 19 instruction-tuned models spanning 0.27B to 32B parameters. Across neutral prompts, richer evidence generally improves evidence-consistent accuracy and ordinal scoring performance. Under user pressure, however, evidence does not reliably prevent user-aligned reversals in this controlled fixed-evidence setting. We report three primary failure modes. First, we identify a negative partial-evidence interaction, where adding epistemic nuance, specifically research gaps, is associated with increased susceptibility to sycophancy in families like Llama-3 and Gemma-3. Second, robustness scales non-monotonically: within some families, certain low-to-mid scale models are especially sensitive to adversarial user pressure. Third, models differ in distributional concentration under conflict: some instruction-tuned models maintain sharply peaked ordinal distributions under pressure, while others are substantially more dispersed; in scale-matched Qwen comparisons, reasoning-distilled variants (DeepSeek-R1-Qwen) exhibit consistently higher dispersion than their instruction-tuned counterparts. These findings suggest that, in a controlled fixed-evidence setting, providing richer in-context evidence alone offers no guarantee against user pressure without explicit training for epistemic integrity.

📄 PDF Abstract BibTeX arXiv:2603.20162

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure

2026-07-02 · Dylan Zongmin Liu arxiv

Personal agents will increasingly negotiate on behalf of users: splitting costs with other personal agents, appealing platform decisions, escalating support disputes, requesting refunds, changing subscriptions, and negot…

PeReGrINE: Evaluating Personalized Review Fidelity with User Item Graph Context

2026-04-09 · Steven Au, Baihan Lin arxiv

We introduce PeReGrINE, a benchmark and evaluation framework for personalized review generation grounded in graph-structured user--item evidence. PeReGrINE restructures Amazon Reviews 2023 into a temporally consistent bi…

Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition

2026-04-07 · Muhammad Ahmed Mohsin, Ahsan Bilal, Muhammad Umer, Emily Fox arxiv

Large language models exhibit sycophancy, the tendency to shift their stated positions toward perceived user preferences or authority cues regardless of evidence. Standard alignment methods fail to correct this because s…

eTracer: Towards Traceable Text Generation via Claim-Level Grounding

2026-01-07 · Bohao Chu, Qianli Wang, Hendrik Damm, Hui Wang 외 arxiv

How can system-generated responses be efficiently verified, especially in the high-stakes biomedical domain? To address this challenge, we introduce eTracer, a plug-and-play framework that enables traceable text generati…

Text Generation

DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models

2026-04-18 · JiYang Wang, Jiawei Chen, Mengqi Xiao, Yu Cheng 외 arxiv

Object level hallucination remains a central reliability challenge for vision language models (VLMs), particularly in binary object existence verification. Existing benchmarks emphasize aggregate accuracy but rarely dise…