OmniFace: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
Abstract
Video Face Swapping (VFS) requires seamlessly injecting asource identity into a target video while meticulously preserving the orig-inal pose, expression, lighting, background, and dynamic information.Existing methods struggle to maintain identity similarity and attributepreservation while preserving temporal consistency. To address the chal-lenge, we propose a comprehensive framework to seamlessly transfer thesuperiority of Image Face Swapping (IFS) to the video domain. Wefirst introduce a novel data pipeline SyncID-Pipe that pre-trains anIdentity-Anchored Video Synthesizer and combines it with IFS modelsto construct bidirectional ID quadruplets for explicit supervision. Build-ing upon paired data, we propose a powerful Diffusion Transformer-basedframework OmniFace, employing a core Modality-Aware Conditioningmodule to discriminatively inject multi-model conditions. Meanwhile,we propose a Synthetic-to-Real Training mechanism and an Identity-Coherence Reward Weighting strategy to enhance visual realism andidentity consistency under challenging scenarios. To address the issue oflimited benchmarks, we introduce IDBench-V, a comprehensive bench-mark encompassing diverse scenes. Extensive experiments demonstrateOmniFace outperforms state-of-the-art methods and further exhibits ex-ceptional versatility, which can be seamlessly adapted to various swap-related tasks.