PhysAlign: Learning Physical Priors for Dynamical Event-Driven Video Generation via Representation Alignment
Abstract
Current video generation models produce highly realistic vi-suals but frequently fail to maintain physical and causal consistencywhen facing complex dynamic events, such as collisions or collapses.We attribute this limitation to the fact that existing models primar-ily fit visual data distributions, entangling dynamics with appearancewithout explicit physical priors to govern motion evolution. To addressthis, we introduce PhysAlign, a framework for Dynamical Event-DrivenPhysically Consistent Video Generation. PhysAlign decouples physicalpriors from visual appearance via a Physical Dynamics Extractor anda discrete Universal Physical Codebook. Recognizing the temporal lo-cality of dynamic events, an Event-Aware Temporal Gating module dy-namically controls the injection of physical priors, while Physics-GuidedAlignment distills continuous physical topologies from a video founda-tion model. Furthermore, we construct EventPhy, a 26K-video datasetwith structured, VLM-generated causal annotations to benchmark dy-namical event-driven video generation. Extensive evaluations show thatPhysAlign achieves state-of-the-art physical reasoning, obtaining leadingSemantic Adherence and Physical Commonsense scores on challengingbenchmarks like PhysicsIQ and VideoPhy, while preserving high visualfidelity. Project Page: https://github.com/FanQi-AI/PhysAlign