OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization
Abstract
Temporal Action Localization (TAL) typically relies on seg-ment annotations or offline access to full videos, limiting scalability andonline use. We introduce Point-Supervised Online TAL (POTAL), whichlocalizes actions in streaming videos using only one temporal point perinstance. To solve POTAL, we propose OnPoint, an offline-to-onlinemulti-level distillation framework that transfers knowledge from a point-supervised offline teacher to an online student via (i) pseudo-segmentinstance distillation, (ii) class-activation sequence distillation, and (iii)anticipatory window-level distillation. We further improve robustness byincorporating the original point labels into student training and by re-fining anchor decoding with actionness-guided attention calibration. Ex-periments on five datasets show OnPoint consistently outperforms strongbaselines, establishing a solid foundation for POTAL. †