paper-with-me

홈 › Papers

Training Alignment Auditors via Reinforcement Learning

2026-08-26 · Paul Rosu, Rowan Wang arxiv

Alignment auditing of frontier models increasingly relies on LLM auditors to surface undesirable behaviors at scale, but current automated auditors can struggle with coherent investigation and audit realism. In this work, we improve LLM auditors with reinforcement learning. In our best training environment, the policy investigates target models that potentially possess hidden behaviors planted via their system prompt. An LLM judge, which knows whether the target has a hidden behavior, holistically compares the policy's investigation to a reference investigation to determine the reward. With systematic ablations, we find that pairwise rewards yield more robust training compared to pointwise rewards, and that adding targets without planted behaviors helps maintain a low false positive rate. Training improves investigation quality against targets with planted behaviors, the rate of concerning behaviors surfaced in unmodified production models, and audit realism, while false-positive rates stay below 1%. Furthermore, auditing capabilities generalize across scaffolds: performance on AuditBench's adversarially fine-tuned targets substantially improves [Sheshadri et al., 2026].

📄 PDF Abstract BibTeX arXiv:2608.25460

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Sequential Auditing for f-Differential Privacy

2026-02-06 · Tim Kutta, Martin Dunsche, Yu Wei, Vassilis Zikas arxiv

We present new auditors to assess Differential Privacy (DP) of an algorithm based on output samples. Such empirical auditors are common to check for algorithmic correctness and implementation bugs. Most existing auditors…

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

2025-10-07 · Matthieu Bou, Nyal Patel, Arjun Jagota, Satyapriya Krishna 외 arxiv

The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a grand challenge. While Inverse Reinforcement Learning (IRL) can infer reward fun…

Reinforcement Learning

Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning

2026-04-24 · Chaoran Chen, Dayu Yuan, Peter Kairouz arxiv

In agentic workflows, LLMs frequently process retrieved contexts that are legally protected from further training. However, auditors currently lack a reliable way to verify if a provider has violated the terms of service…

Reinforcement Learning

Black-Box Access is Insufficient for Rigorous AI Audits

2024-01-25 · Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt 외

External audits of AI systems are increasingly recognized as a key mechanism for AI governance. The effectiveness of an audit, however, depends on the degree of access granted to auditors. Recent audits of state-of-the-a…

Predicting Distresses using Deep Learning of Text Segments in Annual Reports

2018-11-13 · Rastin Matin, Casper Hansen, Christian Hansen, Pia Mølgaard

Corporate distress models typically only employ the numerical financial variables in the firms' annual reports. We develop a model that employs the unstructured textual data in the reports as well, namely the auditors' r…

Descriptive