Conversational Human Audio-visual Talking Dialogue Generation
Abstract
Large-scale dyadic interactive audio-visual dialogue (DIAD)datasets provide fundamental data resources for developing humanoidinteractive virtual agents and digital humans. However, collecting suchdata is time-consuming, expensive, and ethically sensitive. To addressthis, we propose CHAT, a new dyadic interactive audio-visual dialoguegeneration (DIADG) framework that generates diverse, paired, and mu-tually responsive speech-face dialogue clips from a single textual prompt.CHAT unifies large language models and talking face models with inter-active audio and facial behaviour refinement modules, enabling the gen-eration of aligned dyadic dialogue clips with diverse contents and facialidentities. Experiments show that CHAT outperforms existing relatedmethods designed for similar tasks under both objective and subjectiveevaluations. Moreover, our synthesised CHAT-AVD-50k dataset servesas effective pre-training data for downstream interactive head genera-tion, consistently improving PerFRDiff and ReactDiff on REACT 2024.CHAT offers a scalable alternative to the costly and ethically sensitivecollection of real dyadic interaction data.