paper-with-me

홈 › Papers

CounselReflect: A Toolkit for Auditing Mental-Health Dialogues

2026-03-31 · Yahan Li, Chaohao Du, Zeyang Li, Christopher Chun Kuizon, Shupeng Cheng, Angel Hsing-Chi Hwang, Adam C. Frank, Ruishan Liu arxiv

Mental-health support is increasingly mediated by conversational systems (e.g., LLM-based tools), but users often lack structured ways to audit the quality and potential risks of the support they receive. We introduce CounselReflect, an end-to-end toolkit for auditing mental-health support dialogues. Rather than producing a single opaque quality score, CounselReflect provides structured, multi-dimensional reports with session-level summaries, turn-level scores, and evidence-linked excerpts to support transparent inspection. The system integrates two families of evaluation signals: (i) 12 model-based metrics produced by task-specific predictors, and (ii) rubric-based metrics that extend coverage via a literature-derived library (69 metrics) and user-defined custom metrics, operationalized with configurable LLM judges. CounselReflect is available as a web application, browser extension, and command-line interface (CLI), enabling use in real-time settings as well as at scale. Human evaluation includes a user study with 20 participants and an expert review with 6 mental-health professionals, suggesting that CounselReflect supports understandable, usable, and trustworthy auditing. A demo video and full source code are also provided.

📄 PDF Abstract BibTeX arXiv:2603.29429

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FairHealth: An Open-Source Python Library for Trustworthy Healthcare AI in Low-Resource Settings

2026-05-05 · Farjana Yesmin arxiv

We present FairHealth, an open-source Python library that provides a unified, modular framework for trustworthy machine learning in healthcare applications, with particular focus on low-resource and low-income country (L…

Federated Learning

UserSimCRS: A User Simulation Toolkit for Evaluating Conversational Recommender Systems

2023-01-13 · Jafar Afzali, Aleksander Mark Drzewiecki, Krisztian Balog, Shuo Zhang

We present an extensible user simulation toolkit to facilitate automatic evaluation of conversational recommender systems. It builds on an established agenda-based approach and extends it with several novel elements, inc…

Recommendation SystemsText GenerationUser Simulation

MHDash: An Online Platform for Benchmarking Mental Health-Aware AI Assistants

2026-01-30 · Yihe Zhang, Cheyenne N Mohawk, Kaiying Han, Vijay Srinivas Tida 외 arxiv

Large language models (LLMs) are increasingly applied in mental health support systems, where reliable recognition of high-risk states such as suicidal ideation and self-harm is safety-critical. However, existing evaluat…

Dialogue Generation

Auditing Keyword Queries Over Text Documents

2021-12-01 · ICON 2021 12 · Bharath Kumar Reddy Apparreddy, Sailaja Rajanala, Manish Singh

Data security and privacy is an issue of growing importance in the healthcare domain. In this paper, we present an auditing system to detect privacy violations for unstructured text documents such as healthcare records. …

Anomaly Detection

DeepCon: An End-to-End Multilingual Toolkit for Automatic Minuting of Multi-Party Dialogues

2022-09-01 · SIGDIAL (ACL) 2022 9 · Aakash Bhatnagar, Nidhir Bhavsar, Muskaan Singh

In this paper, we present our minuting tool DeepCon, an end-to-end toolkit for minuting the multiparty dialogues of meetings. It provides technological support for (multilingual) communication and collaboration, with a s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationnamed-entity-recognition+6