paper-with-me

홈 › Papers

Democratic ICAI: Debating Our Way to Steering Principles from Preferences

2026-06-26 · Kevin Kingslin, Anish Natekar, Ashutosh Ranjan, Vivek Srivastava, Savita Bhat, Shirish Karande arxiv

Preference-based alignment often struggles to capture the reasoning that underlies human judgments. Many evaluations rely on multiple interacting criteria, yet pairwise labels reveal only the final choice rather than the considerations that shape preferences. Inverse Constitutional AI (ICAI) improves interpretability in decision making by summarizing preferences into natural-language principles, but its single-pass explanations miss much of the nuance involved in complex decisions. We introduce Democratic ICAI, a novel approach that gathers multiple competing rationales through structured persona debate, offering a broader and more expressive account of the factors influencing each comparison. From these richer signals, we derive clearer and more comprehensive steering principles and use them to guide decision modeling through both LLM-based and decision-tree judges. Experiments on creative preference benchmarks, MuCE-Pref and LiTBench, across multiple creative task categories show that Democratic ICAI yields a more faithful preference structure. It improves average preference prediction across tasks relative to deliberative prompting and principle-based baselines, while producing constitutions that LLM annotators prefer.

📄 PDF Abstract BibTeX arXiv:2606.28294

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

Beyond Preferences: Learning Alignment Principles Grounded in Human Reasons and Values

2026-01-26 · Henry Bell, Lara Neubauer da Costa Schertel, Bochu Ding, Brandon Fain arxiv

A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignment models are aligned to a set of princi…

Inverse Constitutional AI: Compressing Preferences into Principles

2024-06-02 · Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier, Samuel Albanie 외

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options, are particularly common. Such preference…

ChatbotLanguage ModellingLarge Language Model

Decoding Human Preferences in Alignment: An Improved Approach to Inverse Constitutional AI

2025-01-28 · Carl-Leander Henneking, Claas Beger

Traditional methods for aligning Large Language Models (LLMs), such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on implicit principles, limiting interpretability. C…

Explainable AI through a Democratic Lens: DhondtXAI for Proportional Feature Importance Using the D'Hondt Method

2024-11-07 · Turker Berk Donmez

In democratic societies, electoral systems play a crucial role in translating public preferences into political representation. Among these, the D'Hondt method is widely used to ensure proportional representation, balanc…

Diabetes PredictionFeature Importance

MentalAgora: A Gateway to Advanced Personalized Care in Mental Health through Multi-Agent Debating and Attribute Control

2024-07-03 · Yeonji Lee, Sangjun Park, Kyunghyun Cho, JinYeong Bak

As mental health issues globally escalate, there is a tremendous need for advanced digital support systems. We introduce MentalAgora, a novel framework employing large language models enhanced by interaction between mult…

AttributeResponse Generation