OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
Abstract
Recent advancements in audio-video joint generation mod-els have demonstrated impressive capabilities in content creation. How-ever, generating high-fidelity human-centric videos in complex, real-worldphysical scenes remains a significant challenge. We identify that the rootcause lies in the structural deficiencies of existing datasets across threedimensions: limited global scene and camera diversity, sparse interac-tion modeling, and insufficient individual attribute alignment. To bridgethese gaps, we present OmniHuman, a large-scale, multi-scene datasetdesigned for fine-grained human modeling. OmniHuman provides a hier-archical annotation covering video-level scenes, frame-level interactions,and individual-level attributes. To facilitate this, we develop a fully au-tomated pipeline for high-quality data collection and multi-modal an-notation. Complementary to the dataset, we establish the OmniHumanBenchmark (OHBench), a three-level evaluation system that providesa scientific diagnosis for human-centric audio-video synthesis. Crucially,OHBench introduces metrics that are highly consistent with human per-ception, filling the gaps in existing benchmarks by providing a compre-hensive diagnosis across global scenes, relational interactions, and in-dividual attributes. Experiments show that fine-tuning on only 20% ofOmniHuman significantly boosts performance, validating its effective-ness in advancing complex scenario modeling. The dataset is available athttps://huggingface.co/datasets/julia527/OmniHuman