LooseControlVideo: Directorial Video Control using Spatial Blocking
Abstract
Precise 3D spatial orchestration in text-to-video generationremains a signix001Ccant challenge, particularly for multi-object scenes wheresemantic layout and temporal dynamics are often entangled. While exist-ing depth-conditioned models achieve good structural x001Cdelity, they neces-sitate dense, frame-accurate guidance that is labor-intensive to authorfor dynamic events involving deformable objects. We present LooseC-ontrolVideo (LCV), a framework that enables intuitive and expressivecontrol by using sparse, oriented 3D boxes as a x0010blockingx0011 proxy. Thisallows users to author high-level layout and trajectory while leveraginga video generative model to generate realistic occlusions, dynamics andinteractions. We achieve this by x001Cne-tuning a Wan 2.2 backbone on avideo dataset annotated with DNOCS, a novel encoding for 3D size, ori-entation and depth-ordered occlusions. Furthermore, our method allowsfor localized rex001Cnementx0016such as adjusting a jump trajectory or addingan interactionx0016with minimal disruption to the global scene context. Ex-tensive evaluations on the nuScenes, HO-3D, and BEHAVE benchmarksdemonstrate that LCV signix001Ccantly outperforms existing 2D-box andx001Dow-based baselines. Our x001Cndings indicate a 1.2-3x improvement in Tra-jectory Error; 2x improvement in Rigid Motion Consistency; and a 1.5-2x increase in Occlusion Accuracy over current state-of-the-art layout-conditioned models, demonstrating that oriented 3D primitives providegood geometric prior for complex, multi-agent video authoring.