World models: predict, then act

If a system can predict what happens after an action, it can compare choices before acting. World-model research asks how to represent environmental change and use predictions for learning and decisions.

A world model learns how an environment changes from observations and actions. If a robot pushes a cup, it may predict where the cup will move; an agent can compare the consequences of several actions. Predictions may concern images or compressed internal states. Their value depends on how they help a task.

Related routes learn useful representations for controllers, train policies on imagined trajectories, or search with a model before acting. Producing plausible videos is another capability. Keep three questions separate when reading: what is predicted, which component is trained, and where the final outcome is evaluated.

Research tasks
  • Compress high-dimensional observations into representations useful for decisions.
  • Generate training trajectories in a learned environment, reducing some real interaction needs.
  • Predict consequences to support search and planning before acting.

Judge a world model by the decisions it supports and whether success inside the model survives evaluation in the actual task.

Predicting what follows an action
  1. 01Observe and encode
  2. 02Predict states
  3. 03Compare futures
  4. 04Act in the world
  5. 05Feedback and update

Research milestones

  1. 2018

    Separate vision, memory, and control

    World Models uses compressed observations and dynamics to support a small controller, including training inside a learned environment.

  2. 2020

    Learn what planning needs

    MuZero learns internal states for reward, value, and policy prediction, then uses tree search. It need not reconstruct future frames.

  3. 2025

    Learn behavior on imagined trajectories

    The formal DreamerV3 paper studies one configuration across task families. Follow how the world model, actor, and critic work together.

Key concepts

Latent state
A compressed internal representation of observations and history. Individual dimensions need not have readable meanings.
Dynamics model
A model predicting how the next state depends on the current state and action.
Imagined rollout
A trajectory generated by repeatedly predicting action outcomes from an internal state. Its feedback comes from the model.
Planning / Policy
Planning compares candidate action sequences at decision time; a policy maps observations or internal states to actions. They can be combined.

World Models and Dreamer

World Models combines compressed observations, dynamics prediction, and control. Follow its diagrams and demos, identifying every module's inputs and outputs.

Then inspect DreamerV3 and how the learned model participates in behavior learning. Relevant foundations include representation learning, sequence prediction, and RL policies and returns.

What to predict, and what to do with predictions

What is predicted? Images, compressed states, or other representations. Representation choices affect training difficulty and the information retained.

What is prediction for? Generating future frames, simulating interaction, training policies, and planning have different requirements. To judge whether actions improve, evaluate downstream task outcomes.

For image and video synthesis, see generation; for robot applications, see embodied AI.

Trace a world-model system

Draw the data and module relationships in World Models. Find a prediction failure and consider its effect on later decisions. Then try running or training under the project's conditions. Long training need not be your first encounter.

Use Awesome World Models by topic. Following one use case builds understanding more readily than treating every system called a world model as the same thing.

World-model papers and research notes

The world-model catalog includes classic projects and paper indexes. Purshow Notes also contains WAM and related technical notes to consult by question.

Representative papers

Begin with the World Models interaction diagram, then compare MuZero and Dreamer: they need different predictions and use them at different times.

2018 · arXiv

World Models

World Models · Ha & Schmidhuber · Fig. 4
Figure excerpt from the paper · Ha & Schmidhuber · Fig. 4 · Original paper

The problem

Can a learned representation support a small controller?

The key idea

Compress images, learn dynamics with a recurrent model, and feed their representations to a controller.

Why this paper

The interactive article makes the components tangible.

Where to start

Read Agent Model and Car Racing, then the imagined VizDoom experiment.

A question to keep asking

Controllers can exploit model errors; test model-trained behavior in the original environment.

2020 · Nature

Mastering Atari, Go, chess and shogi by planning with a learned model

Mastering Atari, Go, chess and shogi by planning with a learned model · Schrittwieser et al. (2020), Fig. 1
Figure excerpt from the paper · Schrittwieser et al. (2020), Fig. 1 · Original paper

The problem

Can a learned model support search without supplied dynamics?

The key idea

Predict reward, value, and policy in internal states, then compare choices with tree search.

Why this paper

Useful models need not reconstruct every pixel.

Where to start

Follow planning, acting, and training before game results.

A question to keep asking

The paper evaluates specific games and action interfaces.

2025 · Nature

Mastering diverse control tasks through world models

Mastering diverse control tasks through world models · Hafner et al. (2025), Fig. 1 (CC BY 4.0)
Figure excerpt from the paper · Hafner et al. (2025), Fig. 1 (CC BY 4.0) · Original paper

The problem

How can one algorithm handle different task signal scales?

The key idea

Learn a world model from experience, train actor and critic on imagined trajectories, and stabilize losses and signal scales.

Why this paper

It connects predictive learning with behavior learning and cross-task evaluation.

Where to start

Start with the training diagram, learning algorithm, and ablations.

A question to keep asking

A shared configuration does not mean one pretrained agent solves every task without training.

Getting started

Stuck on a step? Bring your attempt to the AMA ↗