Nexus-Vid: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
Abstract
Connector-based video unified models have demonstratedstrong capability in instruction-grounded video synthesis, but integratinga large high-fidelity generator into the unified training loop is computa-tionally prohibitive, limiting achievable visual quality. We therefore pro-pose Lumos-Nexus, a training-efficient unified video generation frame-work that facilitates the development of strong reasoning-driven gen-eration capabilities while significantly enhancing visual fidelity. Lumos-Nexus adopts a two-stage design: 1) During training, only a lightweightgenerator is aligned with the understanding block to learn to take inreasoning-driven semantic control. 2) During inference, we introduceUnified Progressive Frequency Bridging (UPFB) to progressively handoff generation to a high-capacity pretrained generator in the shared la-tent space, enabling coarse-to-fine refinement and producing high-fidelityvideos without compromising reasoning quality. To fill the gap in reasoning-driven video generation benchmarks, we introduce VR-Bench, which as-sesses a model’s capability to translate inferred intent into coherent andsemantically aligned video content. Extensive experiments demonstratethat Lumos-Nexus achieves substantial gains in visual realism and tem-poral coherence on VBench, while exhibiting strong reasoning-based gen-erative performance on VR-Bench.