paper-with-me

홈 › Papers

On the Safety of Conversational Models: Taxonomy, Dataset, and Benchmark

2021-10-16 · Findings (ACL) 2022 5 · Hao Sun, Guangxuan Xu, Jiawen Deng, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, Minlie Huang

Dialogue safety problems severely limit the real-world deployment of neural conversational models and have attracted great research interests recently. However, dialogue safety problems remain under-defined and the corresponding dataset is scarce. We propose a taxonomy for dialogue safety specifically designed to capture unsafe behaviors in human-bot dialogue settings, with focuses on context-sensitive unsafety, which is under-explored in prior works. To spur research in this direction, we compile DiaSafety, a dataset with rich context-sensitive unsafe examples. Experiments show that existing safety guarding tools fail severely on our dataset. As a remedy, we train a dialogue safety classifier to provide a strong baseline for context-sensitive dialogue unsafety detection. With our classifier, we perform safety evaluations on popular conversational models and show that existing dialogue systems still exhibit concerning context-sensitive safety problems.

📄 PDF Abstract BibTeX arXiv:2110.08466

Code (1)

thu-coai/diasafety 공식 구현 pytorch

Similar Papers 제목 키워드 기반

On the Safety of Conversational Models: Taxonomy, Dataset, and Benchmark

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Dialogue safety problems severely limit the real-world deployment of neural conversational models and have attracted great research interests recently. However, dialogue safety problems remain under-defined and the corre…

Safety Bench: Identifying Safety-Sensitive Situations for Open-domain Conversational Systems

2021-10-16 · ACL ARR October 2021 10 · Anonymous

The social impact of natural language processing and its applications has received increasing attention. Here, we focus on the problem of safety for end-to-end conversational AI. We survey the problem landscape therein,…

SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems

2022-05-01 · ACL 2022 5 · Emily Dinan, Gavin Abercrombie, A. Bergman, Shannon Spruit 외

The social impact of natural language processing and its applications has received increasing attention. In this position paper, we focus on the problem of safety for end-to-end conversational AI. We survey the problem l…

Position

MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation

2025-08-26 · Ernest Lim, Yajie Vera He, Jared Joselowitz, Kate Preston 외 arxiv

Despite the growing use of large language models (LLMs) in clinical dialogue systems, existing evaluations focus on task completion or fluency, offering little insight into the behavioral and risk management requirements…

Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study

2025-05-21 · Donggeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo Yu

Rapid deployment of vision-language models (VLMs) magnifies safety risks, yet most evaluations rely on artificial images. This study asks: How safe are current VLMs when confronted with meme images that ordinary users sh…

valid