paper-with-me

홈 › Papers

Where Concept Erasure Should Occur: Concept-Layer Alignment in Text-to-Video Diffusion Models

2026-05-25 · Yiwei Xie, Ping Liu, Zheng Zhang arxiv

Text-to-video diffusion transformers encode semantic information unevenly across model depth, which constrains effective concept erasure. We identify a representational bottleneck, termed concept-layer topological alignment, under which target concepts exhibit higher separability at certain representational depths. Outside these depths, concept and non-target signals remain strongly entangled, limiting the effectiveness of depth-specific erasure. This observation reframes concept erasure as the problem of identifying representational depths where concept-non-target separation naturally emerges. Motivated by this structural constraint, we introduce CLEAR, a separability-driven optimization framework for concept erasure that explicitly enforces concept-layer alignment. CLEAR operationalizes this principle by formulating layer selection as an optimization problem over concept-non-target separability, rather than relying on layer-agnostic or heuristic choices. To enable this, we introduce a separability-aware objective that favors layers exhibiting stronger concept-non-target separation. Experiments on large-scale text-to-video diffusion models demonstrate that enforcing concept--layer alignment leads to more precise concept suppression while preserving overall generative quality.

📄 PDF Abstract BibTeX arXiv:2605.25941

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models

2025-03-18 · Yuyang Xue, Edward Moroshko, Feng Chen, Jingyu Sun 외

Text-to-Image diffusion models can produce undesirable content that necessitates concept erasure. However, existing methods struggle with under-erasure, leaving residual traces of targeted concepts, or over-erasure, mist…

MANCE: Manifold Aware Concept Erasure

2026-07-04 · Matan Avitan, Yoav Goldberg, Yanai Elazar hf

Concept erasure aims to remove a target concept from a representation while preserving the other information encoded in it. This is difficult because representations encode many concepts that are often correlated with th…

On Defining Erasure Harms for NLP

2026-06-14 · Yu Lu Liu, Arnav Goel, Jackie Chi Kit Cheung, Alexandra Olteanu 외 arxiv

The deployment of NLP systems has raised concerns about harms they might produce, including representational harms. Recent literature has begun to conceptualize and measure one such harm, the harm of erasure. Nevertheles…

Co-occurring associated retained concepts in Diffusion Unlearning

2026-06-23 · Miso Kim, Georu Lee, Yunji Kim, Hoki Kim 외 arxiv

Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remove not only the target concept, but also benign co-occurring concepts. As illustra…

Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models

2025-10-26 · Lexiang Xiong, Chengyu Liu, Jingwen Ye, Yan Liu 외 arxiv

Concept erasure in text-to-image diffusion models is crucial for mitigating harmful content, yet existing methods often compromise generative quality. We introduce Semantic Surgery, a novel training-free, zero-shot frame…

Text-to-Image Generation