paper-with-me

Papers

Coherence Maximization Improves Pluralistic Alignment

2026-06-02 · Taslim Mahbub, Yiding Pei, Shi Feng arxiv

Aligning AI systems with diverse human values requires value specifications grounded in concrete examples, but generating such examples without extensive human supervision remains an open challenge. We investigate what makes these examples effective, using Internal Coherence Maximization (ICM) -- which infers labels by maximizing their mutual predictability -- to generate persona-specific examples that steer a model toward a target group's values, without human supervision. Across four benchmarks spanning classification, preference, and open-ended generation, ICM-inferred in-context examples match the performance of gold labels. Crucially, coherence matters beyond individual label accuracy: with accuracy held constant, more coherent examples generalize substantially better than incoherent ones. For personas underrepresented in pretraining data, targeted human feedback on the questions where the model is least certain about a persona's values yields better generalization than the same number of labels on arbitrary questions. These results identify coherence as a key design principle for scalable value specification, leveraging the diverse human perspectives already encoded in pretrained language models.

📄 PDF Abstract BibTeX arXiv:2606.03110

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Roadmap to Pluralistic Alignment

2024-02-07 · Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon 외

With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, aligning models to serve pluralistic huma…

Self-Pluralising Culture Alignment for Large Language Models

2024-10-16 · Shaoyang Xu, Yongqi Leng, Linhao Yu, Deyi Xiong

As large language models (LLMs) become increasingly accessible in many countries, it is essential to align them to serve pluralistic human values across cultures. However, pluralistic culture alignment in LLMs remain an …

Prompt Engineering

The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models

2025-07-23 · Giuseppe Russo, Debora Nozza, Paul Röttger, Dirk Hovy arxiv

People increasingly rely on Large Language Models (LLMs) for moral advice, which may influence humans' decisions. Yet, little is known about how closely LLMs align with human moral judgments. To address this, we introduc…

Being Considerate as a Pathway Towards Pluralistic Alignment for Agentic AI

2024-11-15 · Parand A. Alamdari, Toryn Q. Klassen, Rodrigo Toro Icarte, Sheila A. McIlraith

Pluralistic alignment is concerned with ensuring that an AI system's objectives and behaviors are in harmony with the diversity of human values and perspectives. In this paper we study the notion of pluralistic alignment…

Diversity

Pluralistic Off-policy Evaluation and Alignment

2025-09-15 · Chengkai Huang, Junda Wu, Zhouhang Xie, Yu Xia 외 arxiv

Personalized preference alignment for LLMs with diverse human preferences requires evaluation and alignment methods that capture pluralism. Most existing preference alignment datasets are logged under policies that diffe…

Response Generation