paper-with-me

홈 › Papers

T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model

2025-10-25 · Chenyu Zhang, Tairen Zhang, Lanjun Wang, Ruidong Chen, Wenhui Li, Anan Liu arxiv

Using risky text prompts, such as pornography and violent prompts, to test the safety of text-to-image (T2I) models is a critical task. However, existing risky prompt datasets are limited in three key areas: 1) limited risky categories, 2) coarse-grained annotation, and 3) low effectiveness. To address these limitations, we introduce T2I-RiskyPrompt, a comprehensive benchmark designed for evaluating safety-related tasks in T2I models. Specifically, we first develop a hierarchical risk taxonomy, which consists of 6 primary categories and 14 fine-grained subcategories. Building upon this taxonomy, we construct a pipeline to collect and annotate risky prompts. Finally, we obtain 6,432 effective risky prompts, where each prompt is annotated with both hierarchical category labels and detailed risk reasons. Moreover, to facilitate the evaluation, we propose a reason-driven risky image detection method that explicitly aligns the MLLM with safety annotations. Based on T2I-RiskyPrompt, we conduct a comprehensive evaluation of eight T2I models, nine defense methods, five safety filters, and five attack strategies, offering nine key insights into the strengths and limitations of T2I model safety. Finally, we discuss potential applications of T2I-RiskyPrompt across various research fields. The dataset and code are provided in https://github.com/datar001/T2I-RiskyPrompt.

📄 PDF Abstract BibTeX arXiv:2510.22300

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

2024-02-07 · Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu 외

In the rapidly evolving landscape of Large Language Models (LLMs), ensuring robust safety measures is paramount. To meet this crucial need, we propose \emph{SALAD-Bench}, a safety benchmark specifically designed for eval…

DiversityMultiple-choice

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

2024-11-13 · Yangyang Guo, Fangkai Jiao, Liqiang Nie, Mohan Kankanhalli

The vulnerability of Vision Large Language Models (VLLMs) to jailbreak attacks appears as no surprise. However, recent defense mechanisms against these attacks have reached near-saturation performance on benchmark evalua…

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks

2025-05-20 · Guobin Shen, Dongcheng Zhao, Linghao Feng, Xiang He 외

Large language models (LLMs) have achieved remarkable capabilities but remain vulnerable to adversarial prompts known as jailbreaks, which can bypass safety alignment and elicit harmful outputs. Despite growing efforts i…

LLM JailbreakSafety Alignment

Benchmarking Misuse Mitigation Against Covert Adversaries

2025-06-06 · Davis Brown, Mahdi Sabbaghi, Luze Sun, Alexander Robey 외

Existing language model safety evaluations focus on overt attacks and low-stakes tasks. Realistic attackers can subvert current safeguards by requesting help on small, benign-seeming tasks across many independent queries…

BenchmarkingLanguage ModelingLanguage Modelling

Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms

2026-04-26 · Qi Li, Bo Yin, Weiqi Huang, Ruhao Liu 외 arxiv

Vision-Language-Action (VLA) models are emerging as a unified substrate for embodied intelligence. This shift raises a new class of safety challenges, stemming from the embodied nature of VLA systems, including irreversi…