OpenNWM.
A world beyond
the familiar.
Learning to imagine the way ahead — from diverse, action-free videos to controllable navigation.
Motion is everywhere.
NavAnywhere / real video
One observation.
Many steps ahead.
An imagined future, next to the real one. Explore autoregressive rollouts, driven by physical actions.
About these visualizations
Curated qualitative examples from recorded model rollouts; playback is not live inference. Numerical comparisons below use the paper’s benchmark tables.
Different worlds.
A common language.
Streets, homes, gardens, and paths less traveled. Navigation experience across people, robots, and drones.
Explore the dataset
From seeing motion
to understanding it.
A shared latent-action space connects broad visual experience with executable navigation actions.
The latent action model is trained first. World-model training then proceeds through latent pretraining, physical-action adapter warmup, and joint post-training. At inference, planning optimizes real waypoint actions.
Broader experience.
Better prediction.
A 280M-parameter model with improved perceptual prediction on the evaluated navigation benchmarks.
LPIPS, in-domain
DreamSim, in-domain
Reported manuscript results · Table 1. The video examples above are qualitative selections, not a substitute for aggregate evaluation.
The next step
starts here.
A Generalist Navigation World Model with Latent Action Pretraining
The code repository currently requires collaborator access.
@unpublished{opennwm,
title = {OpenNWM: A Generalist Navigation World Model with Latent Action Pretraining},
author = {Anonymous Authors},
note = {Under review as a conference paper at ICLR 2027}
}
Anonymous manuscript · Under review at ICLR 2027. Citation is provisional.