paper-with-me

홈 › Papers

Scaling Trends for Lie Detector Oversight in Preference Learning

2026-07-02 · Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day, Adam Gleave, Chris Cundy arxiv

Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers. In this paper, we scale SOLiD to larger models and evaluate it in more diverse and realistic preference-learning settings. We find favorable scaling: undetected deception drops from 34% for 1B-parameter models to 14% for 405B-parameter models at a detector true positive rate of 99%, and expensive human labelers can be removed entirely from the fine-tuning phase without a statistically significant increase in deception. However, SOLiD is sensitive to distribution shift between detector training and preference-training data, which can drive detector false positive rates to impractical levels.

📄 PDF Abstract BibTeX arXiv:2607.01567

Code (1)

Tavish9/awesome-daily-AI-arxiv ★ 111

Similar Papers 제목 키워드 기반

Preference Learning with Lie Detectors can Induce Honesty or Evasion

2025-05-20 · Chris Cundy, Adam Gleave

As AI systems become more capable, deceptive behaviors can undermine evaluation and mislead users at deployment. Recent work has shown that lie detectors can accurately classify deceptive behavior, but they are not typic…

WorldPM: Scaling Human Preference Modeling

2025-05-15 · Binghai Wang, Runji Lin, Keming Lu, Le Yu 외

Motivated by scaling laws in language modeling that demonstrate how test loss scales as a power law with model and dataset sizes, we find that similar laws exist in preference modeling. We propose World Preference Modeli…

Language ModelingLanguage Modelling

Aligning Object Detector Bounding Boxes with Human Preference

2024-08-20 · Ombretta Strafforello, Osman S. Kayhan, Oana Inel, Klamer Schutte 외

Previous work shows that humans tend to prefer large bounding boxes over small bounding boxes with the same IoU. However, we show here that commonly used object detectors predict large and small boxes equally often. In t…

Object

Inverse Scaling: When Bigger Isn't Better

2023-06-15 · Ian R. McKenzie, Alexander Lyzhov, Michael Pieler, Alicia Parrish 외

Work on scaling laws has found that large language models (LMs) show predictable improvements to overall loss with increased scale (model size, training data, and compute). Here, we present evidence for the claim that LM…

Scaling Laws For Scalable Oversight

2025-04-25 · Joshua Engels, David D. Baek, Subhash Kantamneni, Max Tegmark

Scalable oversight, the process by which weaker AI systems supervise stronger ones, has been proposed as a key strategy to control future superintelligent systems. However, it is still unclear how scalable oversight itse…

Chatbot