paper-with-me

홈 › Papers

PsychBench: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations

2026-04-19 · Patrick Keough arxiv

Large language models are increasingly deployed to simulate patients for clinical training, research, and mental health tools, yet population-level validity remains largely untested. We introduce PsychBench, the first epidemiological audit of LLM patient simulation: 28,800 profiles from four frontier models (GPT-4o-mini, DeepSeek-V3, Gemini-3-Flash, GLM-4.7) evaluated against NHANES and NESARC-III baselines across 120 intersectional cohorts. The central finding is a coherence-fidelity dissociation: models produce clinically plausible individuals while misrepresenting the populations they are drawn from. Variance compression ranges from 14 percent (GLM-4.7) to 62 percent (DeepSeek-V3), eliminating the distributional tails of clinical reality. Despite test-retest correlations above r = 0.90, 36.66 percent of cases cross diagnostic thresholds between runs. Symptom correlation matrices diverge across demographic groups beyond split-half noise, with transgender populations diverging three to five times more than racial differences. Calibration bias is systematic and asymmetric. Models overestimate depression severity for most groups by 3.6 to 6.1 points (Cohen d = 1.13 to 1.91), consistent with training on clinical corpora with elevated base rates. For transgender women the direction inverts: models capture only 8 to 46 percent of documented minority stress elevation, yielding a -5.42 residual (d = -1.55). Models also attribute irritability to Black men and fatigue to women beyond matched controls, encoding racialized and gendered assumptions. Patterns replicate across US and Chinese architectures, indicating failures tied to current training paradigms rather than isolated implementations. For most users, LLM mental health tools risk pathologizing ordinary distress; for transgender users, algorithmic erasure of genuine need. The patients look right. They do not represent real populations.

📄 PDF Abstract BibTeX arXiv:2604.17359

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Generative Approach for Semantic Auditing of Electronic Health Records

2025-07-03 · Irena Girshovitz, Atai Ambus, Moni Shahar, Ran Gilad-Bachrach arxiv

The reliability of clinical artificial intelligence (AI) depends on high-quality data, yet Electronic Health Records are often inconsistent with existing scientific knowledge. Current quality assessments are limited: the…

PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice

2025-02-28 · Shuyu Liu, Ruoxi Wang, Ling Zhang, Xuequan Zhu 외

The advent of Large Language Models (LLMs) offers potential solutions to address problems such as shortage of medical resources and low diagnostic consistency in psychiatric clinical practice. Despite this potential, a r…

BenchmarkingDiagnostic

From Statistical Fidelity to Clinical Consistency: Scalable Generation and Auditing of Synthetic Patient Trajectories

2026-03-06 · Guanglin Zhou, Armin Catic, Motahare Shabestari, Matthew Young 외 arxiv

Access to electronic health records (EHRs) for digital health research is often limited by privacy regulations and institutional barriers. Synthetic EHRs have been proposed as a way to enable safe and sovereign data shar…

Evian: Towards Explainable Visual Instruction-tuning Data Auditing

2026-04-22 · Zimu Jia, Mingjie Xu, Andrew Estornell, Jiaheng Wei arxiv

The efficacy of Large Vision-Language Models (LVLMs) is critically dependent on the quality of their training data, requiring a precise balance between visual fidelity and instruction-following capability. Existing datas…

Logical Fallacies

Auditing and Generating Synthetic Data with Controllable Trust Trade-offs

2023-04-21 · Brian Belgodere, Pierre Dognin, Adam Ivankay, Igor Melnyk 외

Real-world data often exhibits bias, imbalance, and privacy risks. Synthetic datasets have emerged to address these issues. This paradigm relies on generative AI models to generate unbiased, privacy-preserving data while…

Model SelectionPrivacy Preserving