SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning
Abstract
Image-conditioned video diffusion models achieve impressivevisual realism but often suffer from weakened motion fidelity, e.g., re-duced motion dynamics or degraded long-term temporal coherence, es-pecially after fine-tuning. We study motion alignment in video diffusionmodels post-training. To address this, we introduce pixel-motion rewardsbased on pixel flux dynamics, capturing both instantaneous and long-term motion consistency. We further propose Smooth Hybrid Fine-tuning(SHIFT), a scalable reward-driven framework that unifies supervisedfine-tuning and advantage-weighted fine-tuning. Benefiting from noveladversarial advantages, SHIFT improves convergence speed and miti-gates reward hacking. Experiments show that our approach efficientlyresolves dynamic-degree collapse in modern video diffusion models su-pervised fine-tuning. Project page: https://xiye20.github.io/projects/SHIFT/.