paper-with-me

홈 › Papers

LLM Personas as a Substitute for Field Experiments in Method Benchmarking

2025-12-24 · Enoch Hyunwook Kang arxiv

Field experiments (A/B tests) are often the most credible benchmark for methods (algorithms) in societal systems, but their cost and latency bottleneck rapid methodological progress. LLM-based persona simulation offers a cheap synthetic alternative, yet it is unclear whether replacing humans with personas preserves the benchmark interface that adaptive methods optimize against. We prove an if-and-only-if characterization: when (i) methods observe only the aggregate outcome (aggregate-only observation) and (ii) evaluation depends only on the submitted artifact and not on the method's identity or provenance (method-blind evaluation), swapping humans for personas is just panel change from the method's point of view, indistinguishable from changing the evaluation population (e.g., New York to Jakarta). Furthermore, we move from validity to usefulness: we define an information-theoretic discriminability of the induced aggregate channel and show that making persona benchmarking as decision-relevant as a field experiment is fundamentally a sample-size question, yielding explicit bounds on the number of independent persona evaluations required to reliably distinguish meaningfully different methods at a chosen resolution.

📄 PDF Abstract BibTeX arXiv:2512.21080

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Can Digital Personas Reliably Approximate Human Survey Findings?

2026-05-11 · Mumin Jia, Yilin Chen, Divya Sharma, Jairo Diaz-Rodriguez arxiv

Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer t…

Dual Task Framework for Improving Persona-grounded Dialogue Dataset

2022-02-11 · Minju Kim, Beong-woo Kwak, Youngwook Kim, Hong-in Lee 외

This paper introduces a simple yet effective data-centric approach for the task of improving persona-conditioned dialogue agents. Prior model-centric approaches unquestioningly depend on the raw crowdsourced benchmark da…

Benchmarking

Using Large Language Models to Construct Virtual Top Managers: A Method for Organizational Research

2026-01-26 · Antonio Garzon-Vico, Krithika Sharon Komalapati, Arsalan Shahid, Jan Rosier arxiv

This study introduces a methodological framework that uses large language models to create virtual personas of real top managers. Drawing on real CEO communications and Moral Foundations Theory, we construct LLM-based pa…

Are Personalized Stochastic Parrots More Dangerous? Evaluating Persona Biases in Dialogue Systems

2023-10-08 · Yixin Wan, Jieyu Zhao, Aman Chadha, Nanyun Peng 외

Recent advancements in Large Language Models empower them to follow freeform instructions, including imitating generic or specific demographic personas in conversations. We define generic personas to represent demographi…

Benchmarking

HACHIMI: Scalable and Controllable Student Persona Generation via Orchestrated Agents

2026-03-05 · Yilin Jiang, Fei Tan, Xuanyu Yin, Jing Leng 외 arxiv

Student Personas (SPs) are emerging as infrastructure for educational LLMs, yet prior work often relies on ad-hoc prompting or hand-crafted profiles with limited control over educational theory and population distributio…