paper-with-me

Papers

DialogGuard: Multi-Agent Psychosocial Safety Evaluation of Sensitive LLM Responses

2025-12-01 · Han Luo, Guy Laban arxiv

Large language models (LLMs) now mediate many web-based mental-health, crisis, and other emotionally sensitive services, yet their psychosocial safety in these settings remains poorly understood and weakly evaluated. We present DialogGuard, a multi-agent framework for assessing psychosocial risks in LLM-generated responses along five high-severity dimensions: privacy violations, discriminatory behaviour, mental manipulation, psychological harm, and insulting behaviour. DialogGuard can be applied to diverse generative models through four LLM-as-a-judge pipelines, including single-agent scoring, dual-agent correction, multi-agent debate, and stochastic majority voting, grounded in a shared three-level rubric usable by both human annotators and LLM judges. Using PKU-SafeRLHF with human safety annotations, we show that multi-agent mechanisms detect psychosocial risks more accurately than non-LLM baselines and single-agent judging; dual-agent correction and majority voting provide the best trade-off between accuracy, alignment with human ratings, and robustness, while debate attains higher recall but over-flags borderline cases. We release Dialog-Guard as open-source software with a web interface that provides per-dimension risk scores and explainable natural-language rationales. A formative study with 12 practitioners illustrates how it supports prompt design, auditing, and supervision of web-facing applications for vulnerable users.

📄 PDF Abstract BibTeX arXiv:2512.02282

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Hidden Toll of Social Media News: Causal Effects on Psychosocial Wellbeing

2026-01-20 · Olivia Pal, Agam Goyal, Eshwar Chandrasekharan, Koustuv Saha arxiv

News consumption on social media has become ubiquitous, yet how different forms of engagement shape psychosocial outcomes remains unclear. To address this gap, we leveraged a large-scale dataset of ~26M posts and ~45M co…

SynthAgent: A Multi-Agent LLM Framework for Realistic Patient Simulation -- A Case Study in Obesity with Mental Health Comorbidities

2026-02-09 · Arman Aghaee, Sepehr Asgarian, Jouhyun Jeon arxiv

Simulating high-fidelity patients offers a powerful avenue for studying complex diseases while addressing the challenges of fragmented, biased, and privacy-restricted real-world data. In this study, we introduce SynthAge…

Predicting Readiness to Engage in Psychotherapy of People with Chronic Pain Based on their Pain-Related Narratives Saar

2025-06-25 · Saar Draznin Shiran, Boris Boltyansky, Alexandra Zhuravleva, Dmitry Scherbakov 외

Background. Chronic pain afflicts 20 % of the global population. A strictly biomedical mind-set leaves many sufferers chasing somatic cures and has fuelled the opioid crisis. The biopsychosocial model recognises pain sub…

Large Language ModelSensitivitySentenceSpecificity

Ontologia para monitorar a deficiência mental em seus déficts no processamento da informação por declínio cognitivo e evitar agressões psicológicas e físicas em ambientes educacionais com ajuda da I.A*

2024-01-31 · Bruna Araújo de Castro Oliveira

The intention of this article is to propose the use of artificial intelligence to detect through analysis by UFO ontology the emergence of verbal and physical aggression related to psychosocial deficiencies and their pro…

AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models

2025-09-30 · Yixu Wang, Xin Wang, Yang Yao, Xinyuan Li 외 arxiv

The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, existing static benchmarks are ill-equipped to address the dynamic nature of…