paper-with-me

Papers

Evaluation Awareness Scales Predictably in Open-Weights Large Language Models

2025-09-10 · Maheep Chaudhary, Ian Su, Nikhil Hooda, Nishith Shankar, Julia Tan, Kevin Zhu, Ryan Lagasse, Vasu Sharma, Ashwinee Panda arxiv

Large language models (LLMs) can internally distinguish between evaluation and deployment contexts, a behaviour known as \emph{evaluation awareness}. This undermines AI safety evaluations, as models may conceal dangerous capabilities during testing. Prior work demonstrated this in a single $70$B model, but the scaling relationship across model sizes remains unknown. We investigate evaluation awareness across $15$ models scaling from $0.27$B to $70$B parameters from four families using linear probing on steering vector activations. Our results reveal a clear power-law scaling: evaluation awareness increases predictably with model size. This scaling law enables forecasting deceptive behavior in future larger models and guides the design of scale-aware evaluation strategies for AI safety. A link to the implementation of this paper can be found at https://anonymous.4open.science/r/evaluation-awareness-scaling-laws/README.md.

📄 PDF Abstract BibTeX arXiv:2509.13333

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models

2026-04-02 · Sarath Shekkizhar, Romain Cosentino, Adam Earle arxiv

Standard LLM benchmarks evaluate the assistant turn: the model generates a response to an input, a verifier scores correctness, and the analysis ends. This paradigm leaves unmeasured whether the LLM encodes any awareness…

Instruction Following

In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b

2025-09-25 · Nils Durner arxiv

We probe OpenAI's open-weights 20-billion-parameter model gpt-oss-20b to study how sociopragmatic framing, language choice, and instruction hierarchy affect refusal behavior. Across 80 seeded iterations per scenario, we …

Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

2025-03-12 · Zachary Charles, Gabriel Teston, Lucio Dery, Keith Rush 외

As we scale to more massive machine learning models, the frequent synchronization demands inherent in data-parallel approaches create significant slowdowns, posing a critical challenge to further scaling. Recent work dev…

Language ModelingLanguage Modelling

SPA: 3D Spatial-Awareness Enables Effective Embodied Representation

2024-10-10 · Haoyi Zhu, Honghui Yang, Yating Wang, Jiange Yang 외

In this paper, we introduce SPA, a novel representation learning framework that emphasizes the importance of 3D spatial awareness in embodied AI. Our approach leverages differentiable neural rendering on multi-view image…

GPUNeural RenderingRepresentation Learning

MFA-Net: Multi-Scale feature fusion attention network for liver tumor segmentation

2024-05-07 · Yanli Yuan, Bingbing Wang, Chuan Zhang, Jingyi Xu 외

Segmentation of organs of interest in medical CT images is beneficial for diagnosis of diseases. Though recent methods based on Fully Convolutional Neural Networks (F-CNNs) have shown success in many segmentation tasks, …

SegmentationTumor Segmentation