TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection
Abstract
Traditional video highlight detection relies on a narrow, event-centric definition of saliency, which often fails to generalize to uncon-strained personal videos where highlights are heterogeneous and perspective-dependent. To address this, we introduce TRINITY, a multi-perspectivebenchmark that decomposes highlight saliency into three complemen-tary dimensions, Event, Emotion, and Nature, within a unified tempo-ral framework. Leveraging this multi-faceted view, we propose a shared-backbone multi-branch architecture designed for parallel multi-perspectiveprediction via view-specific experts. Comprehensive experiments demon-strate that our method significantly outperforms state-of-the-art base-lines, achieving gains of +7.15/ + 3.62 mAPρ=15%/50% on Mr. HiSumand +10.82 mAP on YouTube Highlights. These results validate thatmulti-perspective modeling provides a more robust and comprehensiveformulation of video saliency, especially for complex real-world scenar-ios. The benchmark and relevant codes will be released upon acceptance.The benchmark is available at https://huggingface.co/datasets/vanilladucky/TRINITY.