ARMS: Anchor–Relational Motion Streaming for Seamless Solo-Social Motion Transitions
Abstract
Generating temporally continuous and socially coherent hu-man motion from text remains a fundamental challenge, particularly inrealistic streams where people act alone, enter interactions, and later dis-engage. Most existing methods generate fixed-length motion clips understatic agent configurations, which makes them brittle to solo–social tran-sitions and unsuitable for incremental generation over long horizons. Wepropose ARMS, an Anchor–Relational Motion Streaming framework thatunifies solo motion and human–human interaction within a single causalgenerative process. ARMS introduces a dynamics-asymmetric represen-tation that decouples per-person temporal evolution from inter-personalignment via a partner-referenced relative-translation term, enablingseamless switching of social coupling without sacrificing long-horizonstability or spatial consistency between agents. On top of a causal la-tent space, a causal relational diffusion model progressively refines mo-tion segment by segment using only past context, capturing both intra-person temporal dependencies and inter-person relations. Mode-awarerelational gating activates or masks cross-agent connections, allowingthe same model to support both solo and interaction generation. Ex-periments show that ARMS improves transition smoothness and socialcoherence compared to interaction-centric baselines, while also achievingcompetitive results on human–human interaction benchmarks.