Back to Blog

World Models Explained: How Genie 3 and Simulated Realities Define the Next Frontier of AI

Published ToolsBear Research Team
World Models Explained: How Genie 3 and Simulated Realities Define the Next Frontier of AI

Type "a rain-slicked Tokyo alley at night, neon signs, puddles" into a video generator and you get a beautiful ten-second clip. Type the same words into Genie 3 and you get somewhere you can walk into — turn left, splash through a puddle, look back at the sign you passed, and find it still there, exactly as you left it. Nothing is pre-built. There is no game engine, no 3D scene, no code. A neural network is predicting every pixel of every frame, 24 times a second, in response to your movements.

This is a world model, and Google DeepMind calls it "a key stepping stone on the path to AGI." This explainer covers what world models are and how they differ from the language models everyone now uses, how Genie evolved from a 2D toy to a real-time simulator, what the technology is genuinely good for, where it still breaks, and why every major lab is now racing to build one.

World Models Explained: Genie 3 and the next frontier beyond chatbots

Executive Summary

  • A world model predicts consequences, not words. Given the current state and an action, it forecasts what happens next — which is what planning, robotics and "imagination" require.
  • Genie 3 (August 2025) was the first real-time interactive world model: 720p at 24 frames per second, worlds that stay consistent for several minutes, with about a minute of memory for what you changed.
  • Project Genie (January 2026) opened the technology to Google AI Ultra subscribers in the US as an experimental prototype for creating, exploring and remixing worlds.
  • The killer application is training agents. An "unlimited curriculum" of simulated environments lets AI agents and robots practise without real-world cost or risk.
  • It is not a video generator. Veo and Sora produce fixed clips from a prompt; Genie produces a controllable stream that reacts to your actions and to mid-simulation text events.
  • Limits are real: horizons of minutes not hours, imperfect physics, limited character control, high compute. Google itself lists these caveats.
  • Everyone is building one: DeepMind (Genie, Veo), OpenAI (Sora's "world simulator" framing), World Labs, NVIDIA (Cosmos), Meta (V-JEPA).

1. What Is a World Model?

A language model predicts the next token; a world model predicts the next state of the world given an action

A large language model is a next-token predictor. It has read a vast amount of text and learned to guess what comes next, which turns out to be a surprisingly general skill. But it has no direct model of physical cause and effect — it knows the word "gravity" appears near the word "falls," not what falling looks like from the inside.

A world model is a next-state predictor. Given where things are now and an action — move forward, push the box, turn the wheel — it predicts the next state of the world. DeepMind's definition: AI systems that "use their understanding of the world to simulate aspects of it, enabling agents to predict both how an environment will evolve and how their actions will affect it."

Why does that matter? Because planning requires simulation. Before you cross a busy road you run a tiny model of the traffic in your head. Before a chess player moves, they imagine the reply. An AI that can only retrieve and describe cannot plan; an AI that can simulate can. This is why Demis Hassabis frames the future of Gemini itself as becoming a world model that can "make plans and imagine new experiences by understanding and simulating aspects of the world, just as the brain does."

2. The Genie Lineage: From 2D Toy to Real-Time Simulator

Timeline of Genie 1, Genie 2, Genie 3 and Project Genie

DeepMind has been building simulated environments for over a decade — training agents on real-time strategy games and open-ended learning long before Genie. The Genie series applied generative modelling to the environments themselves:

  • Genie 1 (2024) — a proof of concept: playable 2D platformer-style worlds generated from images, learned from internet video without action labels. The question was simply "can we do this at all?"
  • Genie 2 (late 2024) — scaled to 3D environments of any kind, generating new worlds for agents to train in. Consistency held for tens of seconds.
  • Genie 3 (August 2025) — the breakthrough: real-time interaction at 24 fps and 720p, worlds that stay consistent "for several minutes," and a new capability called promptable world events. Its researchers describe it plainly: "no underlying game engine, no structure, no code. It's just a neural network that's predicting every single pixel in reaction to inputs from the user and also the past."
  • Project Genie (January 2026) — an experimental web prototype, powered by Genie 3 plus Nano Banana Pro and Gemini, that lets Google AI Ultra subscribers in the US (18+) create, explore and remix interactive worlds from text or images.

3. Genie 3 Under the Hood

Genie 3 specification card: 24 fps, 720p, minutes of consistency, promptable events

Auto-regressive generation

Frames generated one after another, each conditioned on the prompt, past frames and the user's action

Genie 3 builds the world frame by frame, each one conditioned on the text description, the frames that came before, and the user's latest action. This is why DeepMind contrasts it with static 3D capture methods like NeRFs and Gaussian splatting: those reconstruct a fixed scene you can look around; Genie generates the path ahead as you move, so the world can be dynamic and infinite.

The engineering difficulty is memory. As DeepMind notes, "during the auto-regressive generation of each frame, the model has to take into account the previously generated trajectory that grows with time." If you walk away from a wall and come back, the wall must still be there with the same graffiti. Genie 3 sustains this "for several minutes" of exploration, with "memory recalling changes from specific interactions for up to a minute" — you knock something over, walk off, return within a minute, and it's still on the floor.

Promptable world events

A text event such as add a thunderstorm transforms a generated scene mid-simulation

Beyond moving through the world, you can change it with text while it runs — "make it rain," "add a herd of deer," "turn the day to night." DeepMind calls these promptable world events, and they are what turns a pretty simulation into a training tool: you can inject the rare, dangerous or unexpected scenario an agent needs to learn from without waiting for it to happen naturally.

Grounding in reality

Genie can also be "grounded in Street View data from Google Maps," so a generated world can be anchored to a real location — a bridge between simulation and the actual streets a delivery robot will have to navigate.

4. World Model vs. Video Generator

Comparison table of video generation models versus interactive world models

The confusion is understandable — both produce moving images from text, and DeepMind's own Veo models "exhibit a deep understanding of intuitive physics." The distinction is interactivity:

  • A video generator (Veo, Sora) takes a prompt and returns a fixed clip. You can't steer it once it starts. It optimises for a finished piece of content.
  • A world model (Genie 3) takes a prompt and a continuous stream of actions and returns a stream of frames that respond to them. It optimises for consistency under interaction — the property agents need.

DeepMind describes the two as "progress along different capabilities of world simulation." Veo pushes realism and physics; Genie pushes controllability and consistency. Expect them to converge.

5. What World Models Are For

Loop of a world model generating environments in which an agent acts and learns
The "unlimited curriculum": generate a world, let the agent act, observe, learn, generate another.

Training agents at scale

This is the primary motivation. DeepMind's general-purpose agents (its SIMA line and successors) need vast, varied practice. Real environments are slow, expensive and unsafe; hand-built game engines are limited to what humans design. A world model generates an unlimited curriculum of environments — and because the world model "does not know what the goal is," the agent has to genuinely figure things out rather than exploit a scripted level. This is the concrete sense in which DeepMind means "stepping stone to AGI."

Six application areas for world models: robotics, education, games, film, vehicles, planning

Robotics

Robots learn slowly in the physical world and break things while doing it. Simulate a warehouse, a kitchen or a street — grounded in real map data — inject rare events, and a robot policy can rack up years of experience in days.

Education

DeepMind's own example: let students "explore historical eras, like Ancient Rome" — walk the Forum rather than read about it. The same applies to a beating heart, a cell, a distant planet.

Games, film and previsualisation

Project Genie is essentially a world-prototyping tool: sketch a setting in words or an image, walk through it, remix someone else's. For game designers and filmmakers it collapses the gap between concept and explorable space.

Autonomous vehicles and planning

Simulating "what if" — the child stepping out, the sudden downpour — is the core of safety testing for self-driving systems, and promptable world events are exactly that capability. More broadly, any agent that must plan benefits from a scratchpad where it can try actions before committing.

6. Where World Models Still Fall Short

Five current limitations of world models

Google is unusually candid about Project Genie's limits, listing "world realism and character control" as known weaknesses that are "improving." The broader constraints:

  • Horizon. Minutes of consistency is a breakthrough over seconds, but a training episode or a game session lasts hours. Long-horizon memory is the open research problem.
  • Physics fidelity. Intuitive physics is learned from video, not derived from equations. It is impressively plausible and occasionally wrong — a problem when you're training a robot to rely on it.
  • Character and object control. Fine-grained manipulation — precisely grasping, precise text rendering, multi-agent interaction — remains limited.
  • Compute. Predicting every pixel at 24 fps in real time is extraordinarily expensive, which is why access is gated to Ultra subscribers and trusted testers.
  • Evaluation. There is no agreed benchmark for "how good is this world model?" Consistency, controllability and realism trade off against one another.

7. The Competitive Landscape

Organisations building world models: DeepMind, OpenAI, World Labs, NVIDIA, Meta
As publicly described by each organisation.

DeepMind is furthest along in interactive world models with Genie 3, and pairs it with Veo on the video side. OpenAI has framed Sora since its debut as a step toward "world simulators," emphasising physical plausibility in generated video. World Labs, founded by Fei-Fei Li, is building "spatial intelligence" models that turn images into explorable 3D scenes. NVIDIA's Cosmos platform offers world foundation models aimed squarely at robotics and autonomous-vehicle simulation. Meta's V-JEPA line pursues predictive world models that learn abstract representations of what happens next rather than generating pixels — a different bet on the same idea.

The convergence tells you something: the labs disagree on architecture but agree on the destination. Language got AI to fluent; world models are the bet on getting it to competent.

8. Why This Is the Next Frontier

Staircase from perception models to language models to world models to planning agents

Look at the arc. Perception models taught machines to see. Language models taught them to read, write and reason in words. Multimodal models fused the two. World models are the step that adds consequence — the ability to ask "what happens if?" and get a grounded answer. Agents that can plan are built on that. And a universal assistant that can act in your life, the vision behind Project Astra and Gemini, needs exactly this capacity to anticipate.

Genie 3 is the first system to make that idea tangible enough to walk around in. It is early, expensive and imperfect. It is also, quite plainly, a different kind of AI from the chatbots of the last three years — and the labs betting their next decade on it are telling you where they think intelligence comes from.

Key takeaways

  • World models predict the next state given an action — the foundation of planning, robotics and simulation.
  • Genie 3 generates explorable, consistent 3D worlds in real time at 24 fps / 720p with no game engine, and lets you change the world with text as it runs.
  • The main purpose is an unlimited training curriculum for agents; education, games, film and vehicle simulation follow.
  • Current limits — minutes-long horizon, imperfect physics, limited control, heavy compute — are real and acknowledged.
  • DeepMind, OpenAI, World Labs, NVIDIA and Meta are all building toward the same destination by different routes.

Frequently Asked Questions

What is a world model in AI, in one sentence?

An AI system that predicts how an environment will change in response to actions, so it can simulate, plan and let agents practise — as opposed to a language model, which predicts the next word.

What can Genie 3 actually do?

Generate a photorealistic, interactive 3D world from a text description that you can navigate in real time at 24 fps and 720p, remaining consistent for several minutes, with the ability to alter the world mid-simulation using text ("promptable world events").

Can I try Genie 3?

Project Genie, the prototype built on Genie 3, is available to Google AI Ultra subscribers in the United States (18+). The underlying model is otherwise limited to trusted testers; Google says it aims to expand access over time.

How is Genie 3 different from Sora or Veo?

Sora and Veo generate fixed video clips from a prompt. Genie 3 generates a live stream of frames that responds to your movements and to text events while it runs — it is a simulator you interact with, not a video you watch.

Why do researchers call world models a path to AGI?

Because general intelligence requires planning, and planning requires the ability to imagine consequences. A world model provides that internal simulator, and also supplies an unlimited variety of environments in which agents can learn general skills.


Sources: Google DeepMind, "Genie 3: A new frontier for world models" (Aug 5, 2025) and the Genie model page; Google, "Project Genie: AI world model now available for Ultra users in U.S." (Jan 29, 2026); Google DeepMind podcast with Shlomi Fruchter and Jack Parker-Holder; Demis Hassabis, "Our vision for building a universal AI assistant" (May 2025); The Decoder reporting on Genie 3; public descriptions from OpenAI (Sora), World Labs, NVIDIA (Cosmos) and Meta (V-JEPA). Capability figures are as stated by the developers.

TO

ToolsBear Research Team

Research & Editorial

Written by the ToolsBear Research Team team. We test tools, study market trends, and turn complex topics into clear, actionable guides you can use for your next project.