Towards Spatial Supersensing in the Wild
Abstract
Humans can efficiently parse continuous sensory streams,from hours to years, scaffolding an internal world model that groundsspatial reasoning and prediction. To mimic this capacity, spatial super-sensing [59] challenges multimodal models to move beyond linguistic un-derstanding toward true world modeling. However, their benchmark re-lies on synthetic long videos—formed by concatenating random shortclips—and is mostly limited to household scenes, leaving real-world con-tinuity and diversity underexplored. To address the gap, we introduceVSI-Super-Wild, a large-scale spatial supersensing benchmark builtfrom genuinely long, in-the-wild videos across diverse scenarios. Notably,inspired by cognitive studies on how humans structure experience, wesystematically probe the full triad of world state: the agent (observer),objects (scene items), and the environment (places and global layout).* †Equal contribution. Corresponding author.In total, VSI-Super-Wild comprises 442 real-world long-form videosacross 8 scene categories and 6,980 human-verified question-answer pairs.Evaluating multimodal models on VSI-Super-Wild exposes a funda-mental disconnect: despite advances in static image understanding, mod-els consistently fail at tasks that require coherent world-state trackingover time. We characterize how performance degrades with world-statecomplexity and temporal horizon, and diagnose four failure modes: spa-tial collapse, semantic shortcuts, insufficient update, and instance con-fusion. This taxonomy reveals that models lack the mechanisms to bindobjects, agents, and environments into a unified spatial world model—afundamental gap that defines the path forward for spatial supersensing.