paper-with-me

홈 › Papers

Evaluation Awareness Is Not One Capability: Evidence from Open Language Models

2026-06-22 · Nilesh Nayan, Aishwarya Sampath Kumar, Rishiraj Girmal, Shivani Anilkumar, Sankaran Vaidyanathan, David A. Nader Palacio, Reshmi Ghosh, Soundararajan Srinivasan arxiv

Safety benchmarks assume that test-condition behavior predicts deployment behavior, an assumption that fails if models detect evaluation cues and adapt. This opens a gap between benchmark performance and deployment behavior: compliance measured under test conditions becomes an optimistic upper bound that overstates how safely a model behaves once the evaluation harness is removed. We characterize this evaluation awareness through eight experiments across 37 open-weight models and seven families. (i)Detection is moderate and training-driven (24/37 models exceed chance, best AUROC 0.714 vs.0.819 human, with instruction tuning dominating over scale). (ii)Detection shifts safety behavior (hard refusal drops 5.8 percentage points under hypothetical framing, and 21/140 HarmBench framing effects are significant, with compliance rising up to +30 percentage points. (iii)Representations survive behavioral collapse (probes retain AUROC 0.98 under rewrites that drive behavior below chance, and multi-layer steering causally moves three downstream tasks while random controls do not). (iv)These axes are weakly coupled (only 1/15 correlations are significant, the sole robust link being behavioral detection versus framing resistance, $ρ=-0.79$, $p<0.001$). We call this gap the benchmark illusion: because detectability, behavioral manifestation, and controllability vary independently, it is multivariate rather than a single number, so no single awareness score is a reliable proxy for deployment safety.

📄 PDF Abstract BibTeX arXiv:2606.23583

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Security awareness in LLM agents: the NDAI zone case

2026-03-19 · Enrico Bottazzi, Pia Park arxiv

NDAI zones let inventor and investor agents negotiate inside a Trusted Execution Environment (TEE) where any disclosed information is deleted if no deal is reached. This makes full IP disclosure the rational strategy for…

Large Language Models Often Know When They Are Being Evaluated

2025-05-28 · Joe Needham, Giles Edkins, Govind Pimpale, Henning Bartsch 외

If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior during evaluations, leading to less reliable …

MMLUMultiple-choice

Decomposing and Measuring Evaluation Awareness

2026-05-21 · Changling Li, Terry Jingchen Zhang, Jie Zhang, Zhijing Jin 외 arxiv

Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet the field studies it without a shared foundation, conflating properties…

Mechanisms of Introspective Awareness

2026-03-22 · Uzay Macar, Li Yang, Atticus Wang, Peter Wallich 외 arxiv

Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introspective awareness." We investigate the me…

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

2025-10-09 · Nikhil Reddy Varimalla, Yunfei Xu, Arkadiy Saakyan, Meng Fan Wang 외 arxiv

As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness evaluation in VideoLLMs, we introduce Video…