MeanTalker: Efficient and Expressive Speech-Driven 3D Facial Animation via Geometric-Aware Mean Flow
Abstract
Speech-driven 3D facial animation aims to synthesize natural and synchronized facial motions from arbitrary speech audio. Deterministic models struggle with the complex many-to-many mapping, thus yielding over-smoothed motions and lacking expression nuances. In comparison, diffusion-based approaches excel in capturing expressive distributions but suffer from significant latency due to iterative sampling. To this end, we propose MeanTalker, a novel one-step speech-to-motion generative framework built upon Mean Flow. It utilizes a Dual-Time Denoising Transformer (DTDT) for training stability and Geometric-Aware Trajectory Learning (GATL) for precise manifold alignment. Specifically, DTDT introduces an interleaved dual-time encoding to support a progressive curriculum from modeling instantaneous to average velocities, effectively circumventing the instability of direct mean flow training. Furthermore, GATL rectifies the generation trajectory through synergistic dual-space supervision. By applying explicit vertex constraints and projecting latent velocity errors onto the surface manifold, it strictly regularizes the flow direction to guarantee precisely synchronized one-step synthesis. Extensive experiments demonstrate that MeanTalker establishes state-of-the-art lip-sync accuracy while accelerating inference by over 200× compared to iterative diffusion models. Achieving an exceptional Real-Time Factor (RTF) of 0.004, our method effectively bridges the gap between precise motion generation and real-time deployment.