UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation
Abstract
We propose UniDynamics, a diffusion-based framework forfuture 4D dynamic scenes (RGB, depth, and optical flow) generationfrom a single event-RGB pair, without requiring long histories or con-trol priors as in existing methods, while explicitly modeling future mo-tion fields. The core idea is to leverage event streams to offer an al-ternative motion prior for single-RGB extrapolation, and to enforce ge-ometric and motion constraints throughout generation via multimodalmodeling. Specifically, we design an Event Latent Enhancement (ELE)module to align and enhance event latents into diffusion-injectable con-ditioning features, providing robust initial motion priors and reliabletexture/structure cues. We further introduce a Perceptual DynamicsSpace (PDS) embedded in the multi-scale U-Net, which decouples andadaptively interacts depth and flow while continuously feeding back con-straints to appearance features, improving geometric-motion consistencyfor physically plausible and spatiotemporally coherent prediction. Exper-iments on VKitti2 and DSEC demonstrate state-of-the-art performance,producing high-quality, temporally coherent, and 4D-consistent futurepredictions, especially under challenging high-speed motion blur.