EDM:Event-guided Diffusion Model for Video Shadow Detection in Complex Dynamic Scenes
Abstract
Video shadow detection (VSD) aims to robustly identify shadow regions in videos and serves as a key task for improving scene understanding. Although substantial progress has been made in recent years, existing methods still suffer notable performance degradation in challenging scenarios involving fast motion and drastic illumination changes. To address these limitations, we introduce event cameras into the VSD task and propose an Event-guided Diffusion Model (EDM) tailored for complex dynamic environments. Based on the high temporal resolution and asynchronous sensing characteristics of event cameras, we design a Motion Trajectory Attention (MTA) module and a Bidirectional Mask Generation (BMG) module to capture event-driven temporal cues. These cues are further integrated through a Dual-Modal Temporal Guidance (DMTG) module and injected into the diffusion model as conditional guidance to steer the denoising process. By incorporating event streams as a complementary modality, EDM effectively captures fast scene dynamics and abrupt lighting variations, leading to more robust and accurate video shadow detection. Furthermore, to address the lack of event modalities in existing VSD benchmarks, we construct a multimodal extension of the ViSha dataset, termed Event-ViSha, and collect a new real-world multimodal dataset, the Event-based Video Shadow Dataset (EVSD). Extensive experiments demonstrate that our method consistently outperforms state-of-the-art approaches across multiple benchmarks.