paper-with-me

Papers

When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening

2026-05-22 · Jianfeng Zhu, Megan Korhummel, Ruoming Jin, Karin G. Coifman arxiv

As demand for mental health care outpaces clinician-delivered assessment, scalable screening tools are increasingly needed. Large language models (LLMs) may identify psychiatric risk from patient narratives, but their reliability across diagnoses, demographic subgroups, and evidence-use patterns remains uncertain. We introduce a SCID-anchored benchmark of 555 semi-structured experiential interviews paired with diagnostic reference labels for anxiety disorder, major depressive disorder, post-traumatic stress disorder, and any current mental health disorder. Using zero-shot task-specific prompting, we evaluated five state-of-the-art LLMs and examined whether false-negative errors reflected missed psychiatric evidence or differential weighting of symptom, functional-impairment, and protective-context cues. Performance varied across tasks and models, with accuracy ranging from 0.49 to 0.86 and Matthews correlation coefficients from 0.16 to 0.38. GPT-4.1 Mini and GPT-5 Mini showed the most consistent disorder-specific accuracy. Subgroup analyses found higher depression-classification accuracy among male than female participants, no consistent age-related pattern, and modest non-uniform variation across race strata. Evidence-integration analyses showed that false-negative anxiety and PTSD classifications often contained explicit symptom evidence but were accompanied by preserved functioning, coping ability, or social support. Functional-impairment evidence shifted model outputs toward positive classifications, whereas protective-context evidence shifted outputs away. These findings suggest that LLMs may support scalable psychiatric screening, but their tendency to discount symptom evidence in the presence of preserved functioning or protective context requires careful validation before clinical deployment.

📄 PDF Abstract BibTeX arXiv:2605.23148

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multimodal Detection of COVID-19 Symptoms using Deep Learning & Probability-based Weighting of Modes

2021-09-03 · Meysam Effati, Yu-Chen Sun, Hani E. Naguib, Goldie Nejat

The COVID-19 pandemic is one of the most challenging healthcare crises during the 21st century. As the virus continues to spread on a global scale, the majority of efforts have been on the development of vaccines and the…

Machine Learning-driven Analysis of Gastrointestinal Symptoms in Post-COVID-19 Patients

2023-09-27 · Maitham G. Yousif, Fadhil G. Al-Amran, Salman Rawaf, Mohammad Abdulla Grmt

The COVID-19 pandemic, caused by the novel coronavirus SARS-CoV-2, has posed significant health challenges worldwide. While respiratory symptoms have been the primary focus, emerging evidence has highlighted the impact o…

Towards Automatically Classifying Depressive Symptoms from Twitter Data for Population Health

2016-12-01 · WS 2016 12 · Danielle L. Mowery, Albert Park, Craig Bryan, Mike Conway

Major depressive disorder, a debilitating and burdensome disease experienced by individuals worldwide, can be defined by several depressive symptoms (e.g., anhedonia (inability to feel pleasure), depressed mood, difficul…

Artificial Aphasias in Lesioned Language Models

2026-05-15 · Nathan Roll, Jill Kries, Laura Gwilliams, Cory Shain arxiv

Aphasias, selective language impairments which can arise from brain damage, reveal the functional organization of human language by providing causal links between affected brain regions and specific symptom profiles. Dra…

Emergence and dynamics of delusions and hallucinations across stages in early psychosis

2024-02-20 · Catalina Mourgues-Codern, David Benrimoh, Jay Gandhi, Emily A. Farina 외

Hallucinations and delusions are often grouped together within the positive symptoms of psychosis. However, recent evidence suggests they may be driven by distinct computational and neural mechanisms. Examining the time …

Hallucination