Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
Abstract
Motion, speech, and sound effects are fundamental elementsof human-centric videos, yet their heterogeneous temporal characteris-tics make joint generation highly challenging. Existing audio-video gen-eration models often fail to maintain consistent alignment across thesemodalities, leading to noticeable mismatches between motion, speech,and environmental sounds. We present Unison, a unified framework thatexplicitly promotes coherence across the motion, speech, and sound modal-ities. Within the audio stream, Unison employs a semantic-guided har-monization strategy that decouples the generation of speech and sound-effect components. Leveraging bidirectional audio cross-attention andsemantic-conditioned gating for semantic-driven adaptive recomposition,this approach effectively mitigates speech dominance and enhances acous-tic clarity. For audio–motion synchronization, we propose a bidirectionalcross-modal forcing strategy where the cleaner modality guides the nois-ier one through decoupled denoising schedules, reinforced by a progres-sive stabilization strategy. Extensive experiments demonstrate that Uni-son achieves state-of-the-art performance in both audio perceptual qual-ity and cross-modal synchronization, highlighting the importance of ex-plicit multimodal harmonization in human-centric video generation.