paper-with-me

Papers

Understanding Annotator Safety Policy with Interpretability

2026-05-06 · Alex Oesterling, Donghao Ren, Yannick Assogba, Dominik Moritz, Sunnie S. Y. Kim, Leon Gatys, Fred Hohman arxiv

Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from multiple sources such as operational failures (annotators misunderstand or misexecute the task), policy ambiguity (policy wording leaves room for interpretation), or value pluralism (different annotators hold different perspectives on safety). Distinguishing these sources matters. For example, operational failures call for quality control, ambiguity calls for policy clarification, and pluralism calls for deliberation about incorporating diverse perspectives. Yet understanding why annotators disagree is difficult. Directly asking annotators for their reasoning is costly, substantially increasing annotation burden, and can be unreliable for both human and LLM annotators as self-reported reasoning often fails to reflect actual decision processes. We introduce Annotator Policy Models (APMs), interpretable models that learn annotators' internal safety policies from labeling behavior alone, making annotator reasoning visible and comparable without additional annotation effort. We validate that APMs accurately model annotator safety policy (>80% accuracy), faithfully predict responses to counterfactual edits, and recover known policy differences in controlled settings. Applying APMs to LLM and human annotations, we demonstrate two core applications: (1) surfacing policy ambiguity by revealing how annotators interpret safety instructions differently, and (2) surfacing value pluralism by uncovering systematic differences in safety priorities across demographic groups. Together, these capabilities support more targeted, transparent, and inclusive safety policy design.

📄 PDF Abstract BibTeX arXiv:2605.05329

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives

2025-07-21 · Ding Wang, Mark Díaz, Charvi Rastogi, Aida Davani 외 arxiv

Understanding what constitutes safety in AI-generated content is complex. While developers often rely on predefined taxonomies, real-world safety judgments also involve personal, social, and cultural perceptions of harm.…

Position: Require Frontier AI Labs To Release Small "Analog" Models

2025-10-15 · Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer, Philip Quirke arxiv

Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovation tradeoff. This paper argues for an al…

Automated Analysis of Global AI Safety Initiatives: A Taxonomy-Driven LLM Approach

2026-04-04 · Takayuki Semitsu, Naoto Kiribuchi, Kengo Zenitani arxiv

We present an automated crosswalk framework that compares an AI safety policy document pair under a shared taxonomy of activities. Using the activity categories defined in Activity Map on AI Safety as fixed aspects, the …

Mechanistic Interpretability for AI Safety -- A Review

2024-04-22 · Leonard Bereska, Efstratios Gavves

Understanding AI systems' inner workings is critical for ensuring value alignment and safety. This review explores mechanistic interpretability: reverse engineering the computational mechanisms and representations learne…

Large Language Model Safety: A Holistic Survey

2024-12-23 · Dan Shi, Tianhao Shen, Yufei Huang, Zhigen Li 외

The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural language understanding and generation. Howev…

Language ModelingLanguage ModellingLarge Language Modelmodel+2