paper-with-me

Papers

A Roadmap to Pluralistic Alignment

2024-02-07 · Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, Yejin Choi

With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, aligning models to serve pluralistic human values remains an open research question. In this piece, we propose a roadmap to pluralistic alignment, specifically using language models as a test bed. We identify and formalize three possible ways to define and operationalize pluralism in AI systems: 1) Overton pluralistic models that present a spectrum of reasonable responses; 2) Steerably pluralistic models that can steer to reflect certain perspectives; and 3) Distributionally pluralistic models that are well-calibrated to a given population in distribution. We also formalize and discuss three possible classes of pluralistic benchmarks: 1) Multi-objective benchmarks, 2) Trade-off steerable benchmarks, which incentivize models to steer to arbitrary trade-offs, and 3) Jury-pluralistic benchmarks which explicitly model diverse human ratings. We use this framework to argue that current alignment techniques may be fundamentally limited for pluralistic AI; indeed, we highlight empirical evidence, both from our own experiments and from other work, that standard alignment procedures might reduce distributional pluralism in models, motivating the need for further research on pluralistic alignment.

📄 PDF Abstract BibTeX arXiv:2402.05070

Code (1)

jfisher52/ai_pluralistic_alignment 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Being Considerate as a Pathway Towards Pluralistic Alignment for Agentic AI

2024-11-15 · Parand A. Alamdari, Toryn Q. Klassen, Rodrigo Toro Icarte, Sheila A. McIlraith

Pluralistic alignment is concerned with ensuring that an AI system's objectives and behaviors are in harmony with the diversity of human values and perspectives. In this paper we study the notion of pluralistic alignment…

Diversity

Pluralistic Off-policy Evaluation and Alignment

2025-09-15 · Chengkai Huang, Junda Wu, Zhouhang Xie, Yu Xia 외 arxiv

Personalized preference alignment for LLMs with diverse human preferences requires evaluation and alignment methods that capture pluralism. Most existing preference alignment datasets are logged under policies that diffe…

Response Generation

Pluralistic Alignment for Healthcare: A Role-Driven Framework

2025-09-12 · Jiayou Zhong, Anudeex Shetty, Chao Jia, Xuanrui Lin 외 arxiv

As large language models are increasingly deployed in sensitive domains such as healthcare, ensuring their outputs reflect the diverse values and perspectives held across populations is critical. However, existing alignm…

VISPA: Pluralistic Alignment via Automatic Value Selection and Activation

2026-01-19 · Shenyan Zheng, Jiayou Zhong, Anudeex Shetty, Heng Ji 외 arxiv

As large language models are increasingly used in high-stakes domains, it is essential that their outputs reflect not average} human preference, rather range of varying perspectives. Achieving such pluralism, however, re…

VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare

2025-02-19 · Anudeex Shetty, Amin Beheshti, Mark Dras, Usman Naseem

Alignment techniques have become central to ensuring that Large Language Models (LLMs) generate outputs consistent with human values. However, existing alignment paradigms often model an averaged or monolithic preference…

BenchmarkingDiversityMultiple-choice