paper-with-me

홈 › Papers

CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection

2025-10-27 · Grace Byun, Rebecca Lipschutz, Sean T. Minton, Abigail Lott, Jinho D. Choi arxiv

Detecting mental health crisis situations such as suicide ideation, rape, domestic violence, child abuse, and sexual harassment is a critical yet underexplored challenge for language models. When such situations arise during user--model interactions, models must reliably flag them, as failure to do so can have serious consequences. In this work, we introduce CRADLE BENCH, a benchmark for multi-faceted crisis detection. Unlike previous efforts that focus on a limited set of crisis types, our benchmark covers seven types defined in line with clinical standards and is the first to incorporate temporal labels. Our benchmark provides 600 clinician-annotated evaluation examples and 420 development examples, together with a training corpus of around 4K examples automatically labeled using a majority-vote ensemble of multiple language models, which significantly outperforms single-model annotation. We further fine-tune six crisis detection models on subsets defined by consensus and unanimous ensemble agreement, providing complementary models trained under different agreement criteria.

📄 PDF Abstract BibTeX arXiv:2510.23845

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CRADLE: Conversational RTL Design Space Exploration with LLM-based Multi-Agent Systems

2025-08-12 · Lukas Krupp, Maximilian Schöffel, Elias Biehl, Norbert Wehn arxiv

This paper presents CRADLE, a conversational framework for design space exploration of RTL designs using LLM-based multi-agent systems. Unlike existing rigid approaches, CRADLE enables user-guided flows with internal sel…

Cooperating Tools for MWE Lexicon Management and Corpus Annotation

2018-08-01 · COLING 2018 8 · Yuji Matsumoto, Akihiko Kato, Hiroyuki Shindo, Toshio Morita

We present tools for lexicon and corpus management that offer cooperating functionality in corpus annotation. The former, named Cradle, stores a set of words and expressions where multi-word expressions are defined with …

ManagementPOS

Cradle: Empowering Foundation Agents Towards General Computer Control

2024-03-05 · Weihao Tan, Wentao Zhang, Xinrun Xu, Haochong Xia 외

Despite the success in specific scenarios, existing foundation agents still struggle to generalize across various virtual scenarios, mainly due to the dramatically different encapsulations of environments with manually d…

Efficient Exploration

Expert-Level Crisis Detection in Mental Health Conversations

2026-06-09 · Grace Byun, Abigail Lott, Rebecca Lipschutz, Sean T. Minton 외 arxiv

Real-world crisis intervention is inherently conversational, yet existing research largely focuses on static texts.Real-world crisis intervention is inherently conversational, yet existing research largely focuses on sta…

DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning

2026-01-20 · Abdurrahim Yilmaz, Ozan Erdem, Ece Gokyayla, Ayda Acar 외 arxiv

Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesi…

Visual Question Answering