Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion
Abstract
Human motion generation has been widely studied across di-verse input modalities — text, music, and video — and recent efforts haveunified these into single multimodal frameworks. However, while morpho-logical factors such as gender and body shape are known to produce dis-tinct kinematic signatures, no existing unified framework incorporatesthis into generation, treating all subjects as morphologically equiva-lent. We present Odoriko, the first unified multimodal motion generationframework that reflects subject bio-morphological information directly insynthesized motion output. Rather than averaging over subject variation,Odoriko generates motion that is consistent with who is moving, not justwhat they are asked to do — across text, music, and video conditionswithin a single model. When explicit morphological information is un-available, Odoriko additionally recovers subject morphology alongsidemotion, unifying estimation and generation in one framework. Extensiveexperiments across text-to-motion, music-to-dance, and video-to-motionbenchmarks demonstrate that Odoriko matches or exceeds prior special-ized models on standard metrics, while enabling morphology-consistentgeneration that no existing unified framework supports. Project page:https://dsshim0125.github.io/odoriko.github.io/