Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Abstract
Current multimodal rex001Dection mechanisms for long video un-derstanding predominantly rely on closed-loop self-rex001Dection within in-ternal parameters. Lacking objective external evidence, models are fre-quently trapped in blind conx001Cdence and often fail to correct errors.Furthermore, applying reinforcement learning to multi-stage rex001Dectionpipelines introduces severe policy coupling, which is exacerbated by acritical scarcity of dedicated training data. To address these limitations,this work proposes Rex001Dect-R1, the x001Crst Evidence-Driven self-correctionframework for long video understanding. The framework constructs athree-stage pipeline consisting of intuition, verix001Ccation, and arbitration.By dynamically retrieving objective visual evidence to verify initial in-tuitions and autonomously executing multiple temporal searches to re-solve conx001Dicts, it completely breaks the hallucination loop. To over-come policy coupling, we design a stage-decoupled reinforcement learn-ing algorithm named SD-GRPO that independently computes advantagefunctions across dix001Berent reasoning stages. Concurrently, we constructa dataset of 120K samples to bridge the training data gap. Extensiveexperiments on benchmarks such as VideoMME and LongVideoBenchdemonstrate that Rex001Dect-R1 achieves state-of-the-art performance. Ourmethod signix001Ccantly improves the genuine rectix001Ccation rate and enablesauthentic self-correction strictly grounded in objective evidence.