paper-with-me

홈 › Papers

Deep Concept Removal

2023-10-09 · Yegor Klochkov, Jean-Francois Ton, Ruocheng Guo, Yang Liu, Hang Li

We address the problem of concept removal in deep neural networks, aiming to learn representations that do not encode certain specified concepts (e.g., gender etc.) We propose a novel method based on adversarial linear classifiers trained on a concept dataset, which helps to remove the targeted attribute while maintaining model performance. Our approach Deep Concept Removal incorporates adversarial probing classifiers at various layers of the network, effectively addressing concept entanglement and improving out-of-distribution generalization. We also introduce an implicit gradient-based technique to tackle the challenges associated with adversarial training using linear classifiers. We evaluate the ability to remove a concept on a set of popular distributionally robust optimization (DRO) benchmarks with spurious correlations, as well as out-of-distribution (OOD) generalization tasks.

📄 PDF Abstract BibTeX arXiv:2310.05755

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeOut-of-Distribution Generalization

Similar Papers 제목 키워드 기반

Concept Removal for Frontier Image Generative Models

2026-06-24 · Aditya Kumar, Pierre Joly, Adam Dziedzic, Franziska Boenisch arxiv

Image generative models are trained on massive, largely uncurated internet-scale datasets that contain undesirable visual concepts. Efficiently removing such concepts from the model generations without degrading the qual…

Continuous Concepts Removal in Text-to-image Diffusion Models

2024-11-30 · Tingxu Han, Weisong Sun, Yanrong Hu, Chunrong Fang 외

Text-to-image diffusion models have shown an impressive ability to generate high-quality images from input textual descriptions. However, concerns have been raised about the potential for these models to create content t…

Knowledge Distillation

Six-CD: Benchmarking Concept Removals for Benign Text-to-image Diffusion Models

2024-06-21 · Jie Ren, Kangrui Chen, Yingqian Cui, Shenglai Zeng 외

Text-to-image (T2I) diffusion models have shown exceptional capabilities in generating images that closely correspond to textual prompts. However, the advancement of T2I diffusion models presents significant risks, as th…

Benchmarking

Six-CD: Benchmarking Concept Removals for Text-to-image Diffusion Models

2025-01-01 · CVPR 2025 1 · Jie Ren, Kangrui Chen, Yingqian Cui, Shenglai Zeng 외

Text-to-image (T2I) diffusion models have shown exceptional capabilities in generating images that closely correspond to textual prompts. However, the advancement of T2I diffusion models presents significant risks, a…

Benchmarking

Machine Unlearning in Hyperbolic vs. Euclidean Multimodal Contrastive Learning: Adapting Alignment Calibration to MERU

2025-03-19 · Àlex Pujol Vidal, Sergio Escalera, Kamal Nasrollahi, Thomas B. Moeslund

Machine unlearning methods have become increasingly important for selective concept removal in large pre-trained models. While recent work has explored unlearning in Euclidean contrastive vision-language models, the effe…

Contrastive LearningMachine Unlearning