The Path to Reconciling Quality and Safety Alignment in Text-to-Image Generation
Abstract
006 Content safety is a fundamental challenge for text-to-image 006007 (T2I) models, yet prevailing methods enforce a debilitating trade-off 007008 between safety and generation quality. We argue that mitigating this 008009 trade-off hinges on addressing systemic challenges in current T2I safety 009010 alignment across data, methods, and evaluation protocols. To this end, 010011 we introduce a unified framework for synergistic safety alignment. First, 011012 to overcome the flawed data paradigm that provides biased optimiza- 012013 tion signals, we develop LibraAlign-100K, the first large-scale dataset 013014 with dual annotations for safety and quality. Second, to address the my- 014015 opic optimization of existing methods focus solely on safety reward, we 015016 propose Synergistic Preference Optimization (T2I-SPO), a novel align- 016017 ment algorithm that extends the DPO paradigm with a composite re- 017018 ward function that integrates generation safety and quality to holistically 018019 model user preferences. Finally, to overcome the limitations of quality- 019020 agnostic and binary evaluation in current protocols, we introduce the 020021 Unified Alignment Score, a holistic, fine-grained metric that fairly quan- 021022 tifies the balance between safety and generative capability. Extensive 022023 experiments demonstrate that T2I-SPO achieves state-of-the-art safety 023024 alignment against a wide range of NSFW concepts, while better main- 024025 taining the model’s generation quality and general capability. This 025026 paper contains harmful text and image examples. 026