MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer
Abstract
Makeup-transfer applies the reference makeup to the sourceface while preserving the source identity. Despite advances in full-faceediting by diffusion-based methods, strong regional controllability, makeupfidelity, and identity preservation remain challenging. The reasons are(i) pixel-to-attention misalignment that causes spillover into non-targetareas and weakens regional control; (ii) unclear transfer/preservationconcept separation under two-image conditioning, leading to couplingbetween makeup attributes and identity; and (iii) the lack of a high-resolution dataset that is identity-consistent and region-labeled for fine-grained supervision. In this paper, we propose MagicMakeup, a dif-fusion transformer-based framework for region-controllable and high-fidelity makeup transfer, built on spatial constraints and concept dis-entanglement. To enable precise region-specific editing while preservingidentity, we propose Token-Aligned Region Gating, which aligns pixelmasks with attention and applies region-specific logit gating. To clarifythe concepts of transfer and preservation, we further introduce Cross-Modal Perception Guidance, which aligns text and image features to en-hance cross-modal concept perception. We also design a pipeline for thegeneration of 1024 × 1024 data pairs through region-specific makeup re-moval and establish a unified benchmark in synthetic and real settings.Extensive quantitative and qualitative experiments show that Magic-Makeup improves regional controllability, makeup fidelity, and identitypreservation, with strong robustness across styles, races, and poses.