paper-with-me

홈 › Papers

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity

2026-05-07 · Florian A. D. Burnat, Brittany I. Davidson arxiv

Safety benchmarks are routinely treated as evidence about how a language model will behave once deployed, but this inference is fragile if behavior depends on whether a prompt looks like an evaluation. We define evaluation-context divergence as an observable within-item change in behavior induced by framing a fixed task as an evaluation, a live deployment interaction, or a neutral request, and present a paired-prompt protocol that measures it in open-weight LLMs while controlling for paraphrase variation, benchmark familiarity, and judge framing-sensitivity. Across five instruction-tuned checkpoints from four open-weight families plus a matched OLMo-3 base/instruct ablation ($20$ paired items, $840$ generations per checkpoint), we find striking heterogeneity. OLMo-3-Instruct alone is eval-cautious -- evaluation framing raises refusal vs. neutral by $11.8$pp ($p=0.007$) and reduces harmful compliance vs. deployment by $3.6$pp ($p=0.024$, $0/20$ items inverted) -- while Mistral-Small-3.2, Phi-3.5-mini, and Llama-3.1-8B are deployment-cautious}, with marginal eval-vs-deployment refusal effects of $-9$ to $-20$pp. The matched OLMo-3 base also exhibits the deployment-cautious pattern, identifying alignment as the inversion stage; within Llama-3.1, the $70$B model preserves direction with attenuated magnitude, ruling out a simple ``small-model effect that reverses at scale.'' One caveat: the cross-family heterogeneity is judge-dependent. Re-judging with a different-family safety classifier (Llama-Guard-3-8B) preserves the within-OLMo eval-cautious direction but flattens the cross-family contrast, indicating that the two judges operationalize distinct constructs.

📄 PDF Abstract BibTeX arXiv:2605.06327

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers

2021-02-02 · NeurIPS 2021 12 · Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 외

As major progress is made in open-ended text generation, measuring how close machine-generated text is to human language remains a critical open problem. We introduce MAUVE, a comparison measure for open-ended text gener…

Text Generation

Drift No More? Context Equilibria in Multi-Turn LLM Interactions

2025-10-09 · Vardhan Dongre, Ryan A. Rossi, Viet Dac Lai, David Seunghyun Yoon 외 arxiv

Large Language Models (LLMs) excel at single-turn tasks such as instruction following and summarization, yet real-world deployments require sustained multi-turn interactions where user goals and conversational context pe…

Instruction Following

Multiple Token Divergence: Measuring and Steering In-Context Computation Density

2025-12-28 · Vincent Herrmann, Eric Alcaide, Michael Wand, Jürgen Schmidhuber arxiv

Measuring the in-context computational effort of language models is a key challenge, as metrics like next-token loss fail to capture reasoning complexity. Prior methods based on latent state compressibility can be invasi…

Mathematical Reasoning

Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models

2025-08-13 · Igor Halperin arxiv

The proliferation of Large Language Models (LLMs) is challenged by hallucinations, critical failure modes where models generate non-factual, nonsensical or unfaithful text. This paper introduces Semantic Divergence Metri…

Measuring Taiwanese Mandarin Language Understanding

2024-03-29 · Po-Heng Chen, Sijia Cheng, Wei-Lin Chen, Yen-Ting Lin 외

The evaluation of large language models (LLMs) has drawn substantial attention in the field recently. This work focuses on evaluating LLMs in a Chinese context, specifically, for Traditional Chinese which has been largel…