paper-with-me

홈 › Papers

The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks

2026-03-04 · Drake Mullens, Stella Shen arxiv

Do expert personas improve language model performance? The Wharton Generative AI Lab reports that they do not, broadcasting to millions via social media the recommendation that practitioners abandon a technique recommended by Anthropic, Google, and OpenAI. We demonstrate that this null finding was structurally predictable. Five core mechanisms precluded detection before data collection began: baseline contamination elevating the starting point to near-ceiling, system prompt hierarchy subordinating experimental manipulation, impossible expert specifications collapsing to generic competence, format constraints suppressing reasoning processes, and provider exclusion limiting generalizability. Controlled trials correcting these limitations reveal what the original design obscured. To test this, we selected the GPQA Diamond hardest questions to prevent baseline pattern matching, forcing reliance on genuine expert reasoning. On items with valid key answers, expert personas achieve ceiling accuracy. They eliminated all baseline errors through confidence amplification. Furthermore, forensic examination of model divergence identified that half of the hardest GPQA items contain chemically or logically indefensible answers. The model's CoT revealed reasoning away from impossible answers, yielding penalization for accurate chemistry. These findings recontextualize the original null results. Methodologically sound persona research faces measurement constraints imposed by benchmark validity limitations. Answering the persona question requires evaluation infrastructure the field does not yet possess.

📄 PDF Abstract BibTeX arXiv:2603.20225

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy

2025-12-05 · Savir Basil, Ina Shapiro, Dan Shapiro, Ethan Mollick 외 arxiv

This is the fourth in a series of short reports that help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. Here, we ask whether assigning personas to m…

PersonaFlow: Boosting Research Ideation with LLM-Simulated Expert Personas

2024-09-19 · Yiren Liu, Pranav Sharma, Mehul Jitendra Oswal, Haijun Xia 외

Developing novel interdisciplinary research ideas often requires discussions and feedback from experts across different domains. However, obtaining timely inputs is challenging due to the scarce availability of domain ex…

Language ModelingLanguage ModellingLarge Language Modelscientific discovery

FOCUS: Decoupling Expert Personas in LLMs to Enhance Domain Expert Capabilities

2026-08-06 · Guanyu Wang, Zidi Zhang, Xu Chu arxiv

Large Language Models (LLMs) can exhibit diverse personas, and activating expert personas has been shown to improve domain expertise and task accuracy. However, existing persona control methods often suffer from cross-do…

Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance

2025-08-27 · Pedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, Benjamin Roth arxiv

Expert persona prompting -- assigning roles such as expert in math to language models -- is widely used for task improvement. However, prior work shows mixed results on its effectiveness, and does not consider when and w…

Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM

2026-03-19 · Zizhao Hu, Mohammad Rostami, Jesse Thomason arxiv

Persona prompting can steer LLM generation towards a domain-specific tone and pattern. This behavior enables use cases in multi-agent systems where diverse interactions are crucial and human-centered tasks require high-l…