paper-with-me

Papers

EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories

2025-12-19 · Lu Wei, Yuta Nakashima, Noa Garcia arxiv

The widespread adoption of text-to-image (T2I) generation has raised concerns about privacy, bias, and copyright violations. Concept erasure techniques offer a promising solution by selectively removing undesired concepts from pre-trained models without requiring full retraining. However, these methods are often evaluated on a limited set of concepts, relying on overly simplistic and direct prompts. To test the boundaries of concept erasure techniques, and assess whether they truly remove targeted concepts from model representations, we introduce EMMA, a benchmark that evaluates five key dimensions of concept erasure over 13 metrics. EMMA goes beyond standard metrics like image quality and time efficiency, testing robustness under challenging conditions, including indirect descriptions, visually similar non-target concepts, and potential gender and ethnicity bias, providing a socially aware analysis of method behavior. Using EMMA, we analyze five concept erasure methods across five domains (objects, celebrities, art styles, NSFW, and copyright). Our results show that existing methods struggle with implicit prompts (i.e., generating the erased concept when it is indirectly referenced) and visually similar non-target concepts (i.e., failing to generate non-target concepts resembling the erased one), while some amplify gender and ethnicity bias compared to the original model. Code and prompts are available at https://github.com/lobsterlulu/EMMA.

📄 PDF Abstract BibTeX arXiv:2512.17320

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EraseBench: Understanding The Ripple Effects of Concept Erasure Techniques

2025-01-16 · Ibtihel Amara, Ahmed Imtiaz Humayun, Ivana Kajic, Zarana Parekh 외

Concept erasure techniques have recently gained significant attention for their potential to remove unwanted concepts from text-to-image models. While these methods often demonstrate success in controlled scenarios, thei…

Don't Forget Your Embeddings: Robust Knowledge Erasure via Precise Editing of Embeddings

2026-06-02 · Clara Haya Suslik, Or Shafran, Mor Geva arxiv

As language models are increasingly deployed in real-world applications, the ability to erase specific knowledge from them becomes critical for safety and compliance. Prominent methods seek persistent removal by updating…

M-ErasureBench: A Comprehensive Multimodal Evaluation Benchmark for Concept Erasure in Diffusion Models

2025-12-28 · Ju-Hsuan Weng, Jia-Wei Liao, Cheng-Fu Chou, Jun-Cheng Chen arxiv

Text-to-image diffusion models may generate harmful or copyrighted content, motivating research on concept erasure. However, existing approaches primarily focus on erasing concepts from text prompts, overlooking other in…

Image Editing

Comprehensive Assessment and Analysis for NSFW Content Erasure in Text-to-Image Diffusion Models

2025-02-18 · Die Chen, Zhiwen Li, Cen Chen, Xiaodan Li 외

Text-to-image (T2I) diffusion models have gained widespread application across various domains, demonstrating remarkable creative potential. However, the strong generalization capabilities of these models can inadvertent…

Erasing Concepts, Steering Generations: A Comprehensive Survey of Concept Suppression

2025-05-26 · Yiwei Xie, Ping Liu, Zheng Zhang

Text-to-Image (T2I) models have demonstrated impressive capabilities in generating high-quality and diverse visual content from natural language prompts. However, uncontrolled reproduction of sensitive, copyrighted, or h…

Adversarial RobustnessDisentanglementSpecificity