C3ASD: Multi-Level Consistency-Driven Representation Learning for Robust Active Speaker Detection
Abstract
Active Speaker Detection determines whether a visible per-son in a video is speaking at each moment. While recent audio–visualfusion methods perform well on clean data, they degrade under real-world corruptions such as background noise, occlusion, or simultaneousmodality degradation. We attribute this limitation to the absence of ex-plicit consistency constraints that promote robust, semantically alignedrepresentations across modalities. Without such guidance, models tendto learn fragile modality-specific shortcuts that fail under corrupted con-ditions. We propose C 3 ASD, a multi-level consistency-driven frameworkwith three complementary constraints: embedding-level inter-modalityconsistency aligns audio-visual representations during speech; sequence-level intra-modality consistency separates speaking and non-speakingclusters via track-aware contrastive learning; and prediction-level con-sistency stabilizes fusion through knowledge distillation. Extensive ex-periments demonstrate significant improvements under diverse audio, vi-sual and joint corruptions, while maintaining competitive performanceon clean data.