CGCE: Classifier-Guided Concept Erasure in Generative Models
Abstract
Recent advancements in large-scale generative models haveenabled the creation of high-quality images and videos, but have alsoraised significant safety concerns regarding the generation of unsafe con-tent. To mitigate this, concept erasure methods have been developed toremove undesirable concepts from pre-trained models. However, existingmethods remain vulnerable to adversarial attacks that can regeneratethe erased content. Moreover, achieving robust erasure often degradesthe model’s generative quality for safe, unrelated concepts, creating adifficult trade-off between safety and performance. To address this chal-lenge, we introduce Classifier-Guided Concept Erasure (CGCE), an effi-cient plug-and-play framework that provides robust concept erasure fordiverse generative models without altering their original weights. CGCEuses a lightweight classifier operating on text embeddings to first detectand then refine prompts containing undesired concepts. By modifyingonly unsafe embeddings at inference time, our method prevents harm-ful content generation while preserving the model’s original quality onbenign prompts. Extensive experiments show that CGCE achieves state-of-the-art robustness against a wide range of red-teaming attacks. Ourapproach also maintains high generative utility, demonstrating a superiorbalance between safety and performance. We showcase the versatility ofCGCE through its successful application to various modern T2I and T2Vmodels, establishing it as a practical and effective solution for safe gen-erative AI.