paper-with-me

Papers

Benchmarking Overton Pluralism in LLMs

2025-12-01 · Elinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei, Michiel A. Bakker arxiv

We introduce OVERTONBENCH, a novel framework for measuring Overton pluralism in LLMs--the extent to which diverse viewpoints are represented in model outputs. We (i) formalize Overton pluralism as a set coverage metric (OVERTONSCORE), (ii) conduct a large-scale U.S.-representative human study (N = 1208; 60 questions; 8 LLMs), and (iii) develop an automated benchmark that closely reproduces human judgments. On average, models achieve OVERTONSCOREs of 0.35--0.41, with DeepSeek V3 performing best; yet all models remain far below the theoretical maximum of 1.0, revealing substantial headroom for improvement. Because repeated large-scale human studies are costly and slow, scalable evaluation tools are essential for model development. Hence, we propose an automated benchmark that achieves high rank correlation with human judgments ($ρ= 0.88$), providing a practical proxy without replacing human assessment. By turning pluralistic alignment from a normative aim into a measurable benchmark, our work establishes a foundation for systematic progress toward more pluralistic LLMs.

📄 PDF Abstract BibTeX arXiv:2512.01351

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration

2024-06-22 · Shangbin Feng, Taylor Sorensen, YuHan Liu, Jillian Fisher 외

While existing alignment paradigms have been integral in developing large language models (LLMs), LLMs often learn an averaged human preference and struggle to model diverse preferences across cultures, demographics, and…

Overton Pluralistic Reinforcement Learning for Large Language Models

2026-02-24 · Yu Fu, Seongho Son, Ilija Bogunovic arxiv

Existing alignment paradigms remain limited in capturing the pluralistic nature of human values. Overton Pluralism addresses this gap by generating responses with diverse perspectives from a single query. This paper intr…

Natural Language InferenceReinforcement Learning

What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models

2026-05-18 · Payal Chandak, Victoria Alkin, David Wu, Maya Dagan 외 arxiv

Medicine is inherently pluralistic. Principles such as autonomy, beneficence, nonmaleficence, and justice routinely conflict, and such ethical dilemmas often sharply divide reasonable physicians. Good clinical practice n…

From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement

2026-05-14 · Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka arxiv

Pluralistic alignment is typically operationalised as preference aggregation: producing responses that span (Overton), steer toward (Steerable), or proportionally represent (Distributional) diverse human values. We argue…

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation

2026-05-08 · Adib Sakhawat, Tahsin Islam, Takia Farhin, Syed Rifat Raiyan 외 arxiv

The evaluation of Large Language Models (LLMs) faces a critical challenge in construct validity, where fragmented benchmarks and ad hoc metrics frequently conflate method variance, such as prompt sensitivity, with true l…