paper-with-me

홈 › Papers

PERSONA: A Reproducible Testbed for Pluralistic Alignment

2024-07-24 · Louis Castricato, Nathan Lile, Rafael Rafailov, Jan-Philipp Fränken, Chelsea Finn

The rapid advancement of language models (LMs) necessitates robust alignment with diverse user values. However, current preference optimization approaches often fail to capture the plurality of user opinions, instead reinforcing majority viewpoints and marginalizing minority perspectives. We introduce PERSONA, a reproducible test bed designed to evaluate and improve pluralistic alignment of LMs. We procedurally generate diverse user profiles from US census data, resulting in 1,586 synthetic personas with varied demographic and idiosyncratic attributes. We then generate a large-scale evaluation dataset containing 3,868 prompts and 317,200 feedback pairs obtained from our synthetic personas. Leveraging this dataset, we systematically evaluate LM capabilities in role-playing diverse users, verified through human judges, and the establishment of both a benchmark, PERSONA Bench, for pluralistic alignment approaches as well as an extensive dataset to create new and future benchmarks. The full dataset and benchmarks are available here: https://www.synthlabs.ai/research/persona.

📄 PDF Abstract BibTeX arXiv:2407.17387

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Slurry-as-a-Service: A Modest Proposal on Scalable Pluralistic Alignment for Nutrient Optimization

2026-03-02 · Rachel Hong, Yael Eiger, Jevan Hutson, Os Keyes 외 arxiv

Pluralistic alignment has emerged as a promising approach for ensuring that large language models (LLMs) faithfully represent the diversity, nuance, and conflict inherent in human values. In this work, we study a high-st…

Pluralistic Off-policy Evaluation and Alignment

2025-09-15 · Chengkai Huang, Junda Wu, Zhouhang Xie, Yu Xia 외 arxiv

Personalized preference alignment for LLMs with diverse human preferences requires evaluation and alignment methods that capture pluralism. Most existing preference alignment datasets are logged under policies that diffe…

Response Generation

Pluralistic Alignment for Healthcare: A Role-Driven Framework

2025-09-12 · Jiayou Zhong, Anudeex Shetty, Chao Jia, Xuanrui Lin 외 arxiv

As large language models are increasingly deployed in sensitive domains such as healthcare, ensuring their outputs reflect the diverse values and perspectives held across populations is critical. However, existing alignm…

VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare

2025-02-19 · Anudeex Shetty, Amin Beheshti, Mark Dras, Usman Naseem

Alignment techniques have become central to ensuring that Large Language Models (LLMs) generate outputs consistent with human values. However, existing alignment paradigms often model an averaged or monolithic preference…

BenchmarkingDiversityMultiple-choice

EpiPersona: Persona Projection and Episode Coupling for Pluralistic Preference Modeling

2026-03-30 · Yujie Zhang, Weikang Yuan, Zhuoren Jiang, Pengwei Yan arxiv

Pluralistic alignment is essential for adapting large language models (LLMs) to the diverse preferences of individuals and minority groups. However, existing approaches often mix stable personal traits with episode-speci…