arXiv · 2609.33445
Concept Score Relearning: A Unified Cross-Architecture Attack on Concept Erasure
Abstract
Concept erasure aims to suppress undesirable knowledge in text-to-image generative models. However, existing robustness evaluations typically rely on relearning attacks tailored to specific model architectures. We study concept reactivation across two substantially different generative paradigms: noise-prediction U-Nets and flow-matching Transformers. We introduce \textbf{Concept Score Relearning (CSR)}, a unified parameter-level framework that reactivates erased concepts by optimizing each model within its native prediction space. CSR requires no external target-concept image dataset and applies the same concept-directed objective to both U-Net-based Stable Diffusion and Transformer-based FLUX. Experiments across diverse concepts and multiple erasure methods demonstrate consistent concept reactivation across both architectures, highlighting the cross-architecture applicability of CSR and the persistent recoverability of apparently erased concepts. For strict nudity, CSR reaches average ASRs of 50.47\% on FLUX and 40.29\% on Stable Diffusion, consistently ranking first across all evaluated safety settings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hong Xi Tae, Jiaming Zhang, Wenwen He, Xuan Wang, Wei Yang Bryan Lim. 2026-09-27. Concept Score Relearning: A Unified Cross-Architecture Attack on Concept Erasure. https://arxiv.org/abs/2609.33445
Cite the original work for its findings. Save a collection to share your selection of sources.