paper-with-me

홈 › Papers

Linear Probe Accuracy Scales with Model Size and Benefits from Multi-Layer Ensembling

2026-04-15 · Erik Nordby, Tasha Pais, Aviel Parrack arxiv

Linear probes can detect when language models produce outputs they "know" are wrong, a capability relevant to both deception and reward hacking. However, single-layer probes are fragile: the best layer varies across models and tasks, and probes fail entirely on some deception types. We show that combining probes from multiple layers into an ensemble recovers strong performance even where single-layer probes fail, improving AUROC by +29% on Insider Trading and +78% on Harm-Pressure Knowledge. Across 12 models (0.5B--176B parameters), we find probe accuracy improves with scale: ~5% AUROC per 10x parameters (R=0.81). Geometrically, deception directions rotate gradually across layers rather than appearing at one location, explaining both why single-layer probes are brittle and why multi-layer ensembles succeed.

📄 PDF Abstract BibTeX arXiv:2604.13386

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLMs Encode How Difficult Problems Are

2025-10-20 · William Lugoloobi, Chris Russell arxiv

Large language models exhibit a puzzling inconsistency: they solve complex problems yet frequently fail on seemingly simpler ones. We investigate whether LLMs internally encode problem difficulty in a way that aligns wit…

Reinforcement Learning

UniPool: A Globally Shared Expert Pool for Mixture-of-Experts

2026-05-07 · Minbin Huang, Han Shi, Chuanyang Zheng, Yimeng Wu 외 arxiv

Modern Mixture-of-Experts (MoE) architectures allocate expert capacity through a rigid per-layer rule: each transformer layer owns a separate expert set. This convention couples depth scaling with linear expert-parameter…

A single-shot measurement of time-dependent diffusion over sub-millisecond timescales using static field gradient NMR

2020-12-23 · Teddy X. Cai, Nathan H. Williamson, Velencia J. Witherspoon, Rea Ravin 외

Time-dependent diffusion behavior is probed over sub-millisecond timescales in a single shot using an NMR static gradient, time-incremented echo train acquisition (SG-TIETA) framework. The method extends the Carr-Purcell…

Introducing Orthogonal Constraint in Structural Probes

2020-12-30 · ACL 2021 5 · Tomasz Limisiewicz, David Mareček

With the recent success of pre-trained models in NLP, a significant focus was put on interpreting their representations. One of the most prominent approaches is structural probing (Hewitt and Manning, 2019), where a line…

MemorizationPositionSentenceWord Embeddings

Incorporating Kinematic Wave Theory into a Deep Learning Method for High-Resolution Traffic Speed Estimation

2021-02-04 · Bilal Thonnam Thodi, Zaid Saeed Khan, Saif Eddin Jabari, Monica Menendez

We propose a kinematic wave-based Deep Convolutional Neural Network (Deep CNN) to estimate high-resolution traffic speed fields from sparse probe vehicle trajectories. We introduce two key approaches that allow us to inc…