paper-with-me

홈 › Papers

SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems

2022-05-01 · ACL 2022 5 · Emily Dinan, Gavin Abercrombie, A. Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, Verena Rieser

The social impact of natural language processing and its applications has received increasing attention. In this position paper, we focus on the problem of safety for end-to-end conversational AI. We survey the problem landscape therein, introducing a taxonomy of three observed phenomena: the Instigator, Yea-Sayer, and Impostor effects. We then empirically assess the extent to which current tools can measure these effects and current systems display them. We release these tools as part of a “first aid kit” (SafetyKit) to quickly assess apparent safety concerns. Our results show that, while current tools are able to provide an estimate of the relative safety of systems in various settings, they still have several shortcomings. We suggest several future directions and discuss ethical considerations.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Position

Similar Papers 제목 키워드 기반

Monte Carlo Expected Threat (MOCET) Scoring

2025-11-20 · Joseph Kim, Saahith Potluri arxiv

Evaluating and measuring AI Safety Level (ASL) threats are crucial for guiding stakeholders to implement safeguards that keep risks within acceptable limits. ASL-3+ models present a unique risk in their ability to uplift…

Capabilities Ain't All You Need: Measuring Propensities in AI

2026-02-20 · Daniel Romero-Alvarado, Fernando Martínez-Plumed, Lorenzo Pacchiardi, Hugo Save 외 arxiv

AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the tendencies of models to exhibit particular…

Can Large Language Models Make Everyone Happy?

2026-02-11 · Usman Naseem, Gautam Siddharth Kashyap, Ebad Shabbir, Sushant Kumar Ray 외 arxiv

Misalignment in Large Language Models (LLMs) refers to the failure to simultaneously satisfy safety, value, and cultural dimensions, leading to behaviors that diverge from human expectations in real-world settings where …

Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems

2026-02-05 · Ziyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia Tsvetkov arxiv

Language models (LMs) are increasingly used in collaboration: multiple LMs trained by different parties collaborate through routing systems, multi-agent debate, model merging, and more. Critical safety risks remain in th…

The Open Autonomy Safety Case Framework

2024-04-08 · Michael Wagner, Carmen Carlan

A system safety case is a compelling, comprehensible, and valid argument about the satisfaction of the safety goals of a given system operating in a given environment supported by convincing evidence. Since the publicati…

Autonomous Vehiclesvalid