NavWM: A Unified Navigation World Model for Foresight-Driven Planning
Abstract
Conventional visual navigation policies often struggle withmyopic decision-making and mode collapse in complex environments.While world models offer a promising alternative, existing paradigms typ-ically isolate perception, generation, and control, failing to capture theirshared spatio-temporal dynamics. In this paper, we propose NavWM, aunified navigation world model that seamlessly integrates latent worldreasoning, multimodal action prediction, and controllable visual gener-ation. At its core, NavWM leverages latent world tokens to distill geo-metric and semantic priors, endowing the agent with robust structuralunderstanding. To overcome the limitations of deterministic policies, weintroduce an anchor-based multimodal trajectory forecasting frameworkthat generates a diverse action space. This inherent diversity explicitlyempowers the generative world model to act as a robust closed-loop plan-ner, utilizing visual foresight to evaluate and select the optimal path.Extensive experiments across diverse robotics datasets demonstrate thatNavWM significantly advances the state-of-the-art, delivering remark-able improvements in both high-fidelity future state generation and zero-shot navigation success.