paper-with-me

홈 › Papers

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis

2026-03-20 · May Lynn Reese, Markela Zeneli, Mindy Ng, Jacob Haimes, Andreea Damien, Elizabeth Stade arxiv

General-purpose Large Language Models (LLMs) are becoming widely adopted by people for mental health support. Yet emerging evidence suggests there are significant risks associated with high-frequency use, particularly for individuals suffering from psychosis, as LLMs may reinforce delusions and hallucinations. Existing evaluations of LLMs in mental health contexts are limited by a lack of clinical validation and scalability of assessment. To address these issues, this research focuses on psychosis as a critical condition for LLM safety evaluation by (1) developing and validating seven clinician-informed safety criteria, (2) constructing a human-consensus dataset, and (3) testing automated assessment using an LLM as an evaluator (LLM-as-a-Judge) or taking the majority vote of several LLM judges (LLM-as-a-Jury). Results indicate that LLM-as-a-Judge aligns closely with the human consensus (Cohen's $κ_{\text{human} \times \text{gemini}} = 0.75$, $κ_{\text{human} \times \text{qwen}} = 0.68$, $κ_{\text{human} \times \text{kimi}} = 0.56$) and that the best judge slightly outperforms LLM-as-a-Jury (Cohen's $κ_{\text{human} \times \text{jury}} = 0.74$). Overall, these findings have promising implications for clinically grounded, scalable methods in LLM safety evaluations for mental health contexts.

📄 PDF Abstract BibTeX arXiv:2604.02359

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems

2025-12-01 · Xiaochuan Li, Ke Wang, Girija Gouda, Shubham Choudhary 외 arxiv

As Large Language Models (LLMs) become integrated into high-stakes domains, there is a growing need for evaluation methods that are both scalable for real-time deployment and reliable for critical decision-making. While …

Scalable Injury-Risk Screening in Baseball Pitching From Broadcast Video

2026-03-05 · Jerrin Bright, Justin Mende, John Zelek arxiv

Injury prediction in pitching depends on precise biomechanical signals, yet gold-standard measurements come from expensive, stadium-installed multi-camera systems that are unavailable outside professional venues. We pres…

Portable Biomechanics Laboratory: Clinically Accessible Movement Analysis from a Handheld Smartphone

2025-07-11 · J. D. Peiffer, Kunal Shah, Irina Djuraskovic, Shawana Anarwala 외 arxiv

Movement directly reflects neurological and musculoskeletal health, yet objective biomechanical assessment is rarely available in routine care. We introduce Portable Biomechanics Laboratory (PBL), a secure platform for f…

SLMJury: Can Small Language Models Judge as Well as Large Ones?

2026-06-05 · Anish Laddha, Nitesh Pradhan, Gaurav Srivastava arxiv

Large language models (LLMs) are widely used as judges for evaluating model outputs, but their high cost, latency, and opacity limit scalability. We introduce SLMJury, a framework for evaluating small language models (SL…

Domain Generalization

Energy-Based Injury Protection Database: Including Shearing Contact Thresholds for Hand and Finger Using Porcine Surrogates

2026-02-23 · Robin Jeanne Kirschner, Anna Huber, Carina M. Micheler, Dirk Müller 외 arxiv

While robotics research continues to propose strategies for collision avoidance in human-robot interaction, the reality of constrained environments and future humanoid systems makes contact inevitable. To mitigate injury…

Collision Avoidance