CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
Abstract
Multimodal Large Language Models have achieved signif-icant success in offline video understanding, yet their application tostreaming videos is severely limited by the linear explosion of visual to-kens, which often leads to Out-of-Memory (OOM) errors or catastrophicforgetting. Existing visual retention and memory management methodstypically rely on uniform sampling, low-level physical metrics, or pas-sive cache eviction. However, these strategies often lack intrinsic seman-tic awareness, potentially disrupting contextual coherence and blurringtransient yet critical semantic transitions. To address these limitations,we propose CurveStream, a training-free, curvature-aware hierarchicalvisual memory management framework. Our approach is motivated bythe key observation that high-curvature regions along continuous featuretrajectories closely align with critical global semantic transitions. Basedon this geometric insight, CurveStream evaluates real-time semantic in-tensity via a Curvature Score and integrates an online K-Sigma dynamicthreshold to adaptively route frames into clear and blurred memorystates under a strict token budget. Evaluations across diverse tempo-ral scales confirm that this lightweight framework, CurveStream, consis-tently yields absolute performance gains of over 10% (e.g., 10.69% onStreamingBench and 13.58% on OVOBench) over respective baselines,establishing new state-of-the-art results for streaming video perception.