paper-with-me

홈 › Papers

Mitigating Covertly Unsafe Text within Natural Language Systems

2022-10-17 · Alex Mei, Anisha Kabir, Sharon Levy, Melanie Subbiah, Emily Allaway, John Judge, Desmond Patton, Bruce Bimber, Kathleen McKeown, William Yang Wang

An increasingly prevalent problem for intelligent technologies is text safety, as uncontrolled systems may generate recommendations to their users that lead to injury or life-threatening consequences. However, the degree of explicitness of a generated statement that can cause physical harm varies. In this paper, we distinguish types of text that can lead to physical harm and establish one particularly underexplored category: covertly unsafe text. Then, we further break down this category with respect to the system's information and discuss solutions to mitigate the generation of text in each of these subcategories. Ultimately, our work defines the problem of covertly unsafe language that causes physical harm and argues that this subtle yet dangerous issue needs to be prioritized by stakeholders and regulators. We highlight mitigation strategies to inspire future researchers to tackle this challenging problem and help improve safety within smart systems.

📄 PDF Abstract BibTeX arXiv:2210.09306

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification

2025-08-28 · Xiangtao Meng, Yingkai Dong, Ning Yu, Li Wang 외 arxiv

Text-to-image (T2I) generative models have achieved remarkable visual fidelity, yet remain vulnerable to generating unsafe content. Existing safety defenses typically intervene internally within the generative model, but…

Text-to-Image Generation

Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models

2023-05-23 · Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes 외

State-of-the-art Text-to-Image models like Stable Diffusion and DALLE$\cdot$2 are revolutionizing how people generate visual content. At the same time, society has serious concerns about how adversaries can exploit such …

Foveate, Attribute, and Rationalize: Towards Physically Safe and Trustworthy AI

2022-12-19 · Alex Mei, Sharon Levy, William Yang Wang

Users' physical safety is an increasing concern as the market for intelligent systems continues to grow, where unconstrained systems may recommend users dangerous actions that can lead to serious injury. Covertly unsafe …

Attribute

Beware What You Autocomplete: Forensic Attribution of Backdoored Code Completions

2026-07-09 · Anjun Gao, Yueyang Quan, Zhuqing Liu, Minghong Fang arxiv

Large language models have enabled powerful code completion systems that assist developers by predicting subsequent lines of code. However, these models remain vulnerable to backdoor attacks, where malicious fine-tuning …

Code Completion

EvilModel: Hiding Malware Inside of Neural Network Models

2021-07-19 · Zhi Wang, Chaoge Liu, Xiang Cui

Delivering malware covertly and evasively is critical to advanced malware campaigns. In this paper, we present a new method to covertly and evasively deliver malware through a neural network model. Neural network models …