LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
Abstract
Egocentric videos capture rich and diverse human–object in-teractions and have emerged as a fundamental resource for understand-ing human activities related to objects. In this context, Video ReferringExpression Comprehension (Video REC), the task of localizing the tem-poral and spatial extent of a referred object in video frames given anatural language query, plays a key role in linking textual descriptionsto observed objects in untrimmed egocentric recordings. However, exist-ing egocentric Video REC benchmarks primarily focus on short videoclips, where some target object appears densely within frames. Suchsettings do not reflect real-world egocentric recordings, which are long-form, untrimmed, and characterized by sparse object occurrences andcomplex activity transitions. To address this limitation, we introduceLongEgoRefer, a novel and challenging benchmark constructed fromlong-form videos in the Ego4D dataset. LongEgoRefer contains 1,498referring expressions with an average video duration of 45 minutes. Thebenchmark exhibits extreme target sparsity, detailed linguistic descrip-tions, and complex human–object interactions embedded in long, dy-namic egocentric narratives. Consequently, it defines a demanding spatio-temporal grounding problem that requires models to identify both whenan event occurs and where the referred object appears within extendedvideo sequences. We evaluate existing Video REC approaches, includ-ing training-free baselines based on vision–language models combinedwith Grounded SAM2. Extensive experiments show that even advancedbaselines and current state-of-the-art models struggle significantly onLongEgoRefer. These results highlight the intrinsic difficulty of long-form egocentric spatio-temporal grounding and emphasize the need formore robust video understanding models. Our benchmark and code areavailable at https://github.com/shunya-kato/LongEgoRefer.LongEgoRefer0s 4345sA transparent, round bowl containing a bakedvegetable dish is shown being removed from anoven. The bowl, encased in a blue protective cloth,is then placed on the Induction stove. Its contents,a creamy vegetable and mushroom mixture, aresubsequently subjected to a brief external inspection.Human-Object Interaction3067s 3111sRefEgo Long-form video Sparse appearance Linguistic complexity0s 5sThere have white colored hot box in theshelf of the room.Fig. 1: Data comparison of LongEgoRefer and RefEgo. (a) Video durations in LongE-goRefer are orders of magnitude longer than those in RefEgo. (b) The appearance rateis significantly lower and sparser in LongEgoRefer compared to RefEgo. (c) To handlethe complexity of long-form videos, captions in LongEgoRefer are substantially longerand more descriptive than those in RefEgo. Dashed lines indicate the mean value foreach distribution.