TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception
Abstract
We present TriO, a multi-modal unsupervised world modelthat predicts 4D occupancy, obstacle segmentation, flow and LiDAR.In contrast to prior work, TriO utilizes three distinct sensor modali-ties (camera, LiDAR, and RADAR) as both inputs and sources of self-supervision, eliminating the need for additional human annotations. Thanksto its novel supervision, the model is able to segment any occupancyfrom the drivable surface, overcoming the limitations of existing open-set methods in handling long-tail objects. TriO achieves state-of-the-artresults in multiple 3D and 4D tasks, including occupancy, flow, and Li-DAR prediction, as well as zero-shot road obstacle segmentation acrossmultiple datasets such as Argoverse 2, and Spotting the Unexpected.Fig. 1: We present TriO, an unsupervised tri-modal occupancy world model designedto perceive and forecast anything that could be considered path-blocking.