paper-with-me

홈 › Papers

A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI

2026-05-29 · Atahan Karagoz arxiv

Current alignment paradigms for generative artificial intelligence rely predominantly on monolithic benchmarking frameworks that reduce the plurality of human judgment to aggregated statistical baselines, thereby obscuring cultural, demographic, and contextual variability in evaluation. We introduce a state-space constrained emulation framework for AI evaluation that replaces singular assessment functions with a structured manifold of synthetic cognitive profiles representing diverse human perspectives. We show that modern generative architectures can instantiate and maintain these evaluative personas with high consistency, enabling a form of pluralistic, perspective-dependent benchmarking that more closely reflects real-world consensus variability. However, we further analyze the stability of these simulated evaluators under sequential inference and stochastic prompt perturbations, revealing systematic degradation in persona coherence that manifests as state-space drift and semantic inconsistency. These findings suggest that static alignment constraints are insufficient for sustaining robust evaluative behavior over time. Instead, we argue for the necessity of embedding dynamic, viability-driven regulatory mechanisms within generative systems to preserve coherent cognitive emulation. By framing persona-based evaluation as a structured dynamical system over latent representation manifolds, this study provides a foundation for more adaptive, human-aligned, and context-sensitive approaches to AI evaluation.

📄 PDF Abstract BibTeX arXiv:2605.31021

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pluralistic Off-policy Evaluation and Alignment

2025-09-15 · Chengkai Huang, Junda Wu, Zhouhang Xie, Yu Xia 외 arxiv

Personalized preference alignment for LLMs with diverse human preferences requires evaluation and alignment methods that capture pluralism. Most existing preference alignment datasets are logged under policies that diffe…

Response Generation

PERSONA: A Reproducible Testbed for Pluralistic Alignment

2024-07-24 · Louis Castricato, Nathan Lile, Rafael Rafailov, Jan-Philipp Fränken 외

The rapid advancement of language models (LMs) necessitates robust alignment with diverse user values. However, current preference optimization approaches often fail to capture the plurality of user opinions, instead rei…

Pluralistic Alignment for Healthcare: A Role-Driven Framework

2025-09-12 · Jiayou Zhong, Anudeex Shetty, Chao Jia, Xuanrui Lin 외 arxiv

As large language models are increasingly deployed in sensitive domains such as healthcare, ensuring their outputs reflect the diverse values and perspectives held across populations is critical. However, existing alignm…

VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare

2025-02-19 · Anudeex Shetty, Amin Beheshti, Mark Dras, Usman Naseem

Alignment techniques have become central to ensuring that Large Language Models (LLMs) generate outputs consistent with human values. However, existing alignment paradigms often model an averaged or monolithic preference…

BenchmarkingDiversityMultiple-choice

A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models

2025-04-09 · Zhouhang Xie, Junda Wu, Yiran Shen, Yu Xia 외

Personalized preference alignment for large language models (LLMs), the process of tailoring LLMs to individual users' preferences, is an emerging research direction spanning the area of NLP and personalization. In this …