paper-with-me

홈 › Papers

AI Alignment at Your Discretion

2025-02-10 · Maarten Buyl, Hadi Khalaf, Claudio Mayrink Verdun, Lucas Monteiro Paes, Caio C. Vieira Machado, Flavio du Pin Calmon

In AI alignment, extensive latitude must be granted to annotators, either human or algorithmic, to judge which model outputs are better' or safer.' We refer to this latitude as alignment discretion. Such discretion remains largely unexamined, posing two risks: (i) annotators may use their power of discretion arbitrarily, and (ii) models may fail to mimic this discretion. To study this phenomenon, we draw on legal concepts of discretion that structure how decision-making authority is conferred and exercised, particularly in cases where principles conflict or their application is unclear or irrelevant. Extended to AI alignment, discretion is required when alignment principles and rules are (inevitably) conflicting or indecisive. We present a set of metrics to systematically analyze when and how discretion in AI alignment is exercised, such that both risks (i) and (ii) can be observed. Moreover, we distinguish between human and algorithmic discretion and analyze the discrepancy between them. By measuring both human and algorithmic discretion over safety alignment datasets, we reveal layers of discretion in the alignment process that were previously unaccounted for. Furthermore, we demonstrate how algorithms trained on these datasets develop their own forms of discretion in interpreting and applying these principles, which challenges the purpose of having any principles at all. Our paper presents the first step towards formalizing this core gap in current alignment processes, and we call on the community to further scrutinize and control alignment discretion.

📄 PDF Abstract BibTeX arXiv:2502.10441

Code (0)

등록된 구현이 없습니다.

Tasks

Safety Alignment

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Discretionary vs nondiscretionary in fiscal mechanism. Non-automatic fiscal stabilisers vs automatic fiscal stabilisers

2025-01-15 · Vasile Bratian, Amelia Bucur, Camelia Oprean, Cristina Tanasescu

The goal of the present study is to increase the intelligibility of macroeconomic phenomena triggered by governmental intervention in economy by means of fiscal policies. During cyclical movements, fiscal policy can play…

Discretionary Trees: Understanding Street-Level Bureaucracy via Machine Learning

2023-12-17 · Gaurab Pokharel, Sanmay Das, Patrick J. Fowler

Street-level bureaucrats interact directly with people on behalf of government agencies to perform a wide range of functions, including, for example, administering social services and policing. A key feature of street-le…

Understanding the Learning Dynamics of Alignment with Human Feedback

2024-03-27 · Shawn Im, Yixuan Li

Aligning large language models (LLMs) with human intentions has become a critical task for safely deploying models in real-world systems. While existing alignment approaches have seen empirical success, theoretically und…

The Politics of (No) Compromise: Information Acquisition, Policy Discretion, and Reputation

2021-10-31 · Liqun Liu

Precise information is essential for making good policies, especially those regarding reform decisions. However, decision-makers may hesitate to gather such information if certain decisions could have negative impacts on…

Decision Making

Discretion in the Loop: Human Expertise in Algorithm-Assisted College Advising

2025-05-19 · Sofiia Druchyna, Kara Schechtman, Benjamin Brandon, Jenise Stafford 외

In higher education, many institutions use algorithmic alerts to flag at-risk students and deliver advising at scale. While much research has focused on evaluating algorithmic predictions, relatively little is known abou…

Heterogeneous Treatment Effect Estimation