LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
Abstract
The evaluation of long-term video quality understanding re-mains an open challenge for large vision-language models (LVLMs). Ex-isting video quality benchmarks predominantly focus on short clips andisolated distortions, overlooking the temporal continuity, cumulative degra-dation, and reasoning complexity inherent in long-duration content. Toaddress these limitations, we present LongVQUBench, a comprehen-sive benchmark for long-term video quality understanding. LongVQUBenchcontains over 1,200 diverse videos spanning movies, documentaries, surveil-lance footage, egocentric recordings, and animated content, accompa-nied by 1,500 multiple-choice and open-ended questions for validationand testing. To assess perceptual reasoning across different temporalscopes, we introduce three progressively complex evaluation levels: (i)local event quality understanding (LQU) for analyzing localized distor-tions; (ii) cross-event quality reasoning (CQR) for integrating multipledegraded events; and (iii) global quality understanding (GQU) for holis-tic perceptual evaluation over extended durations. Furthermore, a needledistortion question-answering (NDQA) paradigm is embedded across allthree levels, where spatial or temporal artifacts are sparsely inserted toprobe fine-grained detection and reasoning capabilities. Extensive ex-periments on 14 state-of-the-art LVLMs reveal significant performancedegradation with increasing video length and reasoning depth, highlight-ing their limited capacity for long-range temporal integration and percep-tual attribution. We envision LongVQUBench as a foundational step to-ward the systematic, hierarchical, and explainable evaluation of LVLMs’long-term video quality understanding.