ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding
Abstract
Spatio-Temporal Video Grounding (STVG) aims to retrievethe visual trajectory of a specific object from a video stream as describedby a natural language expression. However, most advanced methodsstruggle to balance global context modeling with precise boundary lo-calization. Due to the prohibitive computational costs of processing longvideos, these approaches typically resort to low-rate temporal downsam-pling and implicit motion modeling. This inevitably suppresses high-frequency boundary cues and neglects the explicit inter-frame depen-dencies required for precise boundary delineation. To address these lim-itations, we present ScanFocus, a novel coarse-to-fine framework thatdecouples the STVG task into a global spatio-temporal scan and a localboundary focus. Specifically, we utilize a unified vision-language fusionencoder combined with a lightweight Deformable Semantic-Motion Fu-sion module to efficiently align multimodal features and generate coarseproposals. To recover the suppressed fine-grained details, we introducethe Semantic-Guided Temporal Aggregator (SGTA) in the refinementstage. By densely sampling around coarse boundaries, SGTA explicitlymodels short-term temporal interactions under semantic guidance, cap-turing rapid motion changes for precise timestamp regression. Extensiveexperiments on three widely used benchmarks demonstrate the perfor-mance superiority of our proposed method over previous approaches.Code will be released at https://github.com/TenMinutes209/ScanFocus.