CustomX: Unified Character, Action, and Scene Customization in Video World Models
Abstract
Recent advances in world models have greatly enhanced in-teractive environment simulation. Existing methods mainly fall into two⋆ Equal contribution. † Corresponding author.categories: (1) static world generation models, which construct 3D en-vironments without active agents, and (2) controllable-entity models,which allow a single entity to perform limited actions in an otherwise un-controllable environment. In this work, we introduce CustomX, leverag-ing the realism and structural grounding of static world generation whileextending controllable-entity models to support user-specified characterscapable of performing open-ended actions. Users can provide a 3DGSscene and a character, then use natural language to direct the characterto perform diverse behaviors, ranging from basic locomotion to object-centric interactions, while freely exploring the environment. CustomXsynthesizes temporally coherent video clips that preserve visual fidelitywith the provided scene and character, formulated as a conditional au-toregressive video generation problem. Built upon a pre-trained videogenerator, our training strategy significantly enhances motion dynamicswhile maintaining generalization across actions and characters. Our eval-uation covers a broad range of aspects, including visual quality, characterconsistency, action controllability, and long-horizon coherence.