Degradation-Robust and Temporally Consistent Infrared–Visible Video Fusion via One-step Diffusion Framework
Abstract
Infrared–visible video fusion aims to integrate complemen-tary thermal and visual information to produce stable and informativefused videos for real-world applications. While recent advances haveachieved remarkable progress in image-level fusion, extending these meth-ods to dynamic video scenarios remains challenging. Existing video fu-sion approaches typically adopt a ‘temporal-first’ processing strategy,where temporal modeling is performed within each modality before cross-modal fusion. However, such designs may become vulnerable to modality-specific degradations, including low illumination in visible videos andnoise or flicker in infrared sequences, which often lead to unreliablemotion estimation and temporally inconsistent fusion results. To ad-dress this limitation, we advocate a ‘fusion-first’ strategy and proposeDRT-VF, a unified degradation-robust and temporally consistent in-frared–visible video fusion method built upon a one-step diffusion ar-chitecture with a two-stage progressive training strategy. In Stage I, aModality-Collaborative Attention module integrates complementary in-frared and visible cues to establish robust intra-frame fusion representa-tions. In Stage II, temporal consistency is modeled through a hierarchicaldesign consisting of a Global Temporal Consistency module and a LocalMotion Compensation module, which jointly enhance global temporalconsistency and refine motion-sensitive regions. Extensive experimentsdemonstrate that DRT-VF improves temporal stability while preservingfine structural details.