Improved Immiscible Diffusion: Accelerating Diffusion Training by Reducing Miscibility
Abstract
The substantial training cost of di!usion models hinderstheir deployment. Immiscible Di!usion [20] showed that reducing themixing of images’ di!usion destination in the noise space via linear as-signment can accelerate di!usion training. However, concerns regardinglimited image diversity and exploding execution times under large batchsizes limit its feasibility for large-scale training. In this work, we startfrom thoroughly tackling these limitations: For the image diversity, wedemonstrate the bijective nature of the denoising process of vanilla dif-fusion, underlying that being immiscible cannot hurt the diversity. Mov-ing beyond the noise layer, we refine immiscible di!usion’s concept to abroader miscibility reduction at any layer, which enables us to proposea new family of its implementations much more e"cient to execute un-der high batch sizes, including K-nearest neighbor (KNN) noise selectionand image scaling. Overall, the immiscible di!usion family achieves upto 4→ faster training across diverse models and tasks, including uncon-ditional/conditional generation, image editing, and robotics planning.Extensive analysis shows step-by-step on how immiscibility eases denois-ing and improves e"ciency. Besides, our analysis of immiscibility o!ersa novel perspective on how optimal transport (OT) enhances di!usiontraining. By identifying trajectory miscibility as a fundamental bottle-neck, we believe this work establishes a potentially new direction forfuture research in high-e"ciency di!usion training.