paper-with-me

홈 › Papers

An Auditing Test To Detect Behavioral Shift in Language Models

2024-10-25 · Leo Richter, Xuanli He, Pasquale Minervini, Matt J. Kusner

As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task performance, and alignment with societal values. Extensive initial evaluations, including red teaming and diverse benchmarking, can establish a model's behavioral profile. However, subsequent fine-tuning or deployment modifications may alter these behaviors in unintended ways. We present a method for continual Behavioral Shift Auditing (BSA) in LMs. Building on recent work in hypothesis testing, our auditing test detects behavioral shifts solely through model generations. Our test compares model generations from a baseline model to those of the model under scrutiny and provides theoretical guarantees for change detection while controlling false positives. The test features a configurable tolerance parameter that adjusts sensitivity to behavioral changes for different use cases. We evaluate our approach using two case studies: monitoring changes in (a) toxicity and (b) translation performance. We find that the test is able to detect meaningful changes in behavior distributions using just hundreds of examples.

📄 PDF Abstract BibTeX arXiv:2410.19406

Code (1)

richterleo/Auditing_Test_for_LMs 공식 구현 pytorch

Tasks

BenchmarkingChange DetectionRed Teaming

Similar Papers 제목 키워드 기반

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards

2025-11-18 · Yule Liu, Heyi Zhang, Jinyi Zheng, Zhen Sun 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has become a core training stage in recent large language models (LLMs). Its reliance on non-public, high-value prompt sets raises concerns about unauthorized data us…

Reinforcement Learning

Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning

2026-04-24 · Chaoran Chen, Dayu Yuan, Peter Kairouz arxiv

In agentic workflows, LLMs frequently process retrieved contexts that are legally protected from further training. However, auditors currently lack a reliable way to verify if a provider has violated the terms of service…

Reinforcement Learning

Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models

2026-05-06 · Quintin Pope, Ajay Hayagreeve Balaji, Jacques Thibodeau, Xiaoli Fern arxiv

We present an automated, contrastive evaluation pipeline for auditing the behavioral impact of interventions on large language models. Given a base model $M_1$ and an intervention model $M_2$, our method compares their f…

knowledge editing

Offscript: Automated Auditing of Instruction Adherence in LLMs

2025-12-11 · Nicholas Clark, Ryan Bai, Tanu Mitra arxiv

Large Language Models (LLMs) and generative search systems are increasingly used for information seeking by diverse populations with varying preferences for knowledge sourcing and presentation. While users can customize …

Instruction Following

Martingale Doppelgänger-Eval: An Identification Framework for Auditing Candlestick Understanding in Vision-Language Models

2026-06-16 · Ziyao Wang arxiv

We introduce Martingale Doppelgänger-Eval, a public shadow-market benchmark for auditing whether vision-language models (VLMs) use candlestick evidence rather than extrapolate past trends. The central difficulty is ident…