SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset
Abstract
Recent co-speech gesture generation methods often overlookcultural differences, limiting their effectiveness in human–agent inter-action. Moreover, culture-conditioned models are rarely evaluated un-der speaker-disjoint splits, so apparent “cultural” behavior may be con-founded with speaker-specific gesturing style. We introduce SICAGE, amodular framework for culture-aware co-speech gesture generation thatconditions motion synthesis models on speaker-independent cultural rep-resentations. SICAGE learns these representations from audio and textby treating each speaker as a separate domain while imposing invari-ance across speakers. This encourages representations to remain culture-discriminative while reducing dependence on speaker identity. The re-sulting cultural embeddings condition a multimodal generator to pro-duce culturally appropriate gestures. We instantiate this idea with twodomain generalization approaches: adversarial learning and Fishr reg-ularization. We further introduce ALaDiT, a real-time diffusion-basedgesture generator designed to efficiently incorporate the learned culturalembeddings. To validate our method, we built TED4C-L, a 106-hourmultimodal dataset of 764 TED speakers from four cultural groups. Ex-periments show that SICAGE improves motion realism, diversity, beatsynchronization, semantic relevance, and cultural consistency.