World models, from the ground up

"World model" has become one of the most overloaded terms in AI. Startups use it in pitch decks, chip companies sell "world foundation model" platforms, and research labs use it to mean several quite different things. This post builds a precise mental model from the ground up: what a world model is, how it differs from the two things it's most often confused with, how frontier labs actually train one, how anyone measures whether it gets physics right, and who is building them today.

A primer on the learned simulators behind the next wave of robotics, driving, and interactive media. Facts about companies and models are current as of October 2026. Technical terms are collected in a glossary, and sources are listed at the end.

Part One

Foundations

What a world model is, and how it differs from the two things it's most often confused with.

What a world model is

A world model is a learned system that answers one question:

"If the world is in this state, and this action happens, what does the world look like next?"

Three ingredients are involved:

  • State: what the world looks like right now, represented either as images or as a compressed internal summary.
  • Action: something that changes the world, such as a keypress, a steering angle, or a robot arm motion.
  • Prediction: the next state.

Chain those predictions together, feeding each prediction back in as the new state, and you get a simulator you can "play" or plan inside. The crucial difference from a traditional simulator is that this one is learned from data rather than written by hand.

A chess engine can look ahead because it knows exactly how each move changes the board. A world model tries to learn the "rules" of a messy physical world (gravity, occlusion, how a door swings, how a pedestrian behaves) purely by watching enough of it, so an agent can look ahead the same way.

Two camps under one name

The term covers two families that are worth keeping separate:

  • Generative world models predict actual images, video, or 3D scenes you can look at. Examples: Google DeepMind's Genie, Odyssey, World Labs, NVIDIA Cosmos.
  • Predictive world models (sometimes called latent world models) predict only an abstract summary of the future and never render pixels, because their purpose is planning rather than viewing. Generative models work in a compressed internal space too, as the training section shows; the difference is that predictive models never turn it back into pictures. Examples: Yann LeCun's JEPA line of research (now at AMI Labs) and Dreamer-style reinforcement learning agents.

Most of this post focuses on the generative camp, since that's where most of the current activity and funding sits, with notes where the predictive camp differs. For an intuitive take on the debate between the two camps, with a worked example of a ball rolling behind a box, see my earlier post, Demystifying AI models in robotics.

That definition hinges on one word: action. It's also exactly what separates a world model from the thing it's most often mistaken for.

How world models differ from video generators like Sora

The defining difference is action-conditioning: the ability to change what happens next based on an input that arrives mid-stream.

A video generator like Sora takes a prompt and produces a complete clip. You are a spectator. A world model generates the future step by step in response to inputs as they arrive. You are a participant. When DeepMind released the first Genie in 2024, its research lead made this exact point: a video generator can be visually stunning, but a world model needs actions.

That one requirement forces three others:

  1. It must run in real time, because it's waiting on your next input. Genie 3 generates at 24 frames per second in 720p.
  2. It must generate step by step (autoregressively, explained in Autoregressive vs. diffusion), building each frame from the previous ones rather than producing the whole clip at once.
  3. It must remember. If you turn away from a building and turn back, the building should still be there. A 10-second clip never has to worry about this.

The two are close cousins, though. Video models are often the starting point for world models: Genie 3 builds on both its predecessor Genie 2 and DeepMind's Veo 3 video model. OpenAI made the relationship explicit in March 2026 when it shut down the Sora app and API, stating that Sora's original goal had been to teach AI to understand and simulate the physical world, and that the Sora research team would continue on world simulation research aimed at robotics.

A video generator is, in effect, a world model without the actions. A physics engine is the opposite case: all hand-written rules, nothing learned.

How world models differ from game physics engines

Game engines like Unity rely on a classical physics engine (Unity's 3D physics runs on NVIDIA PhysX). It's worth being precise about what that is:

  • It stores an explicit list of objects with known positions, masses, and shapes.
  • It advances them using hand-written equations of motion and collision rules.
  • It does not draw anything. A separate renderer turns the state into pixels, using 3D assets a human artist built.

A world model collapses all three jobs (the state, the physics, and the rendering) into one learned network. That produces a clean set of trade-offs:

Physics engineWorld model
AccuracyExact and deterministicApproximate; can hallucinate
GuaranteesConservation laws hold by constructionNo hard guarantees; objects can morph
ContentOnly what someone authoredNew worlds from a sentence or photo
Hard phenomena cloth, fluids, crowdsDifficult to write equations forLearned from data, often convincingly
Visual realismLimited by authored assetsPhotorealism comes "for free" from data
Cost to runCheapRoughly a GPU per user

Google itself has been careful about this distinction, describing its Project Genie as not a game engine and not capable of a full game experience.

The physics engine is an accountant applying the tax code line by line. The world model is a veteran tax preparer who has seen a million returns and can tell you what the answer looks like: usually right, occasionally confidently wrong.

Part Two

How they work

Part One defined what a world model does. This part opens it up: the generation technique underneath, how frontier labs train one, and where quality comes from.

Autoregressive vs. diffusion, and why world models need both

Two generation techniques sit at the heart of every modern generative world model. They're often presented as rivals, but they actually answer different questions.

The two ideas

  • Autoregressive generation answers "how do you build a long sequence?" One piece at a time, each piece conditioned only on what came before, never revising what's already been produced. This is how a large language model writes text, token by token.
  • Diffusion answers "how do you build one complex sample?" Start from pure random noise and refine all of it in parallel over many small cleanup steps, from coarse to fine, until a sharp image emerges.

Autoregression is a novelist writing chapter by chapter, unable to go back and edit chapter 2 once chapter 3 is written. Diffusion is a sculptor working on a whole block at once, roughing out the shape first and then refining every surface together.

Why neither works alone

Pure diffusion over a whole clip is roughly how Sora-class video generators work. All frames are denoised simultaneously, and every frame can "see" every other frame during generation (called bidirectional attention), so the future can influence the past. That gives excellent coherence within a clip, but it's fatal for a world model: the clip length is fixed in advance, and you can't inject an action at frame 40 because frame 40 was already being shaped alongside frame 100 from the very first step. Looking ahead is fine for a movie but impossible in an interactive world where the next input hasn't happened yet.

Pure autoregression over discrete tokens was the earlier world-model approach (the original Genie and Wayve's GAIA-1 worked roughly this way). Each frame is chopped into a grid of codes from a fixed vocabulary, LLM-style, and the model predicts the next frame's codes. This is naturally interactive, but compressing images into a fixed vocabulary throws away detail, and the results tend to be blurrier and glitchier than diffusion.

The hybrid: autoregressive across time, diffusion within each frame

Modern world models (Genie 3, Odyssey-3, World Labs' Atlas, NVIDIA Cosmos) combine the two on separate axes. Both Atlas and Odyssey-3 describe themselves as autoregressive diffusion transformers, and this is what that phrase means.

Diffusion within one frame (vertical); autoregression across time (horizontal).

The loop at every tick:

  1. Take the clean frames generated so far, plus the action that just arrived.
  2. Start the next frame as pure noise.
  3. Denoise it in a handful of steps while "looking at" that context.
  4. Append the finished frame to the context and repeat.

The model's attention is causal in time: a frame can see the past but never the future. The past is cached so it isn't recomputed every tick, exactly like an LLM's KV cache. Many systems generate a small chunk of frames per tick instead of one, trading a little latency for smoother motion.

What makes the hybrid hard

Exposure bias. During training, the model is normally given perfect, real past frames as context. At inference, it gets its own slightly imperfect outputs instead. It has never practiced recovering from its own mistakes, so small errors accumulate into drift. This is the root of the compounding error problem discussed under open challenges. Several techniques target it:

  • Noise augmentation: deliberately corrupt context frames during training so the model learns not to trust them blindly. Google's GameNGen (the neural Doom simulator) relied on this.
  • Diffusion Forcing (Chen et al., 2024): train each frame with its own independent noise level, teaching the model to treat the past as reliable and the future as uncertain on a sliding scale.
  • Self Forcing (2025): train the model on its own generated rollouts, so practice matches deployment.

Speed. At 24 frames per second there are about 40 milliseconds per frame, and diffusion normally takes dozens of denoising steps. The standard fix is distillation: train a slow, high-quality "teacher" model first, then train a fast "student" to reproduce its output in one to four steps. A widely copied recipe, CausVid, distills a bidirectional teacher (which can look ahead) into a causal few-step student, keeping the teacher's quality while gaining real-time interactivity.

Memory. The KV cache can't grow forever, so models use sliding windows, compressed summaries of older frames, or explicit 3D memory structures (see frontier research directions).

A note on the predictive camp: JEPA-style models are autoregressive in time but use no diffusion at all, because they never generate pixels. They predict the next abstract representation directly.

That loop is what a finished world model runs. Getting a model that can run it takes a training pipeline of six stages.

How frontier world models are trained

At the architecture level, a generative world model has three jobs: compress video into something manageable, learn dynamics (how the compressed world evolves under actions), and make it fast enough for real time. The actual training pipeline spreads those jobs across six stages, much as an LLM pipeline has pretraining, instruction tuning, and reinforcement learning from feedback.

Curate data
↓
Train tokenizer Compress
↓
Pretrain on video Learn dynamicsmillions of hours
↓
Action mid-training Learn dynamicstens of thousands of hours
↓
Distill for real time Make fast
↓
Post-train
CompressLearn dynamicsMake fast

Where the data comes from: the data pyramid

The single most important fact about world-model data is that it forms a pyramid. There is an enormous amount of video with no actions attached, and only a tiny amount where you know exactly which action caused what happened.

NVIDIA's published numbers make this concrete. The original Cosmos was trained on 20 million hours of raw video, filtered down to 100 million clips. The action-labeled set used for Cosmos 3 contains 8.4 million episodes totaling about 61,300 hours: roughly 0.3% of the raw pool. Every lab's data strategy is essentially an answer to "how do we make the top of the pyramid bigger?"

From most abundant to scarcest:

  • Internet video. Volume and visual diversity (every material, lighting condition, and place on earth), but no actions. It's also biased toward edited, interesting footage with camera cuts and very few failures.
  • Egocentric human video. Head-mounted camera footage of people cooking, assembling, and walking. Head and hand motion can be estimated from the footage and treated as actions. In Cosmos 3's action data, egocentric motion is the largest component at about 41,300 hours (67%).
  • Video games. The classic action-labeled source, since every controller input can be recorded alongside every frame. The first Genie was trained on internet footage of 2D platformer games.
  • Driving logs. Camera feeds paired with exact steering, braking, and speed. This is why self-driving companies (Waymo, Wayve, Tesla) hold a structural data advantage, and why Odyssey's founders came from that world.
  • Robot data. The scarcest and most valuable. In Cosmos 3, robotics contributes only about 5,400 hours (9%), aggregated from open-source datasets.
  • Purpose-captured data. Footage collected specifically for training. Odyssey, for example, built its own 360-degree camera backpack rig early on to film real places.
  • Synthetic data from physics engines. Perfectly labeled ground truth, at the cost of looking like a simulator.

The camera-motion trick. Any video shot from a moving camera implicitly contains a "where did the viewer move?" signal, recoverable with 3D reconstruction tools. For navigation-style world models (walk forward, turn left), that's a free action label on millions of hours of footage. Cosmos 3 mined about 4,600 hours of camera-motion data from its own pretraining videos this way.

Stage by stage

1. Curate data. Raw video is mostly unusable: talking heads, slideshows, static shots, watermarks, duplicates. NVIDIA's pipeline is representative: split long videos into individual shots, filter out low-value clips, write a text caption for each clip, remove near-duplicates, and group clips by resolution and aspect ratio. Captioning matters more than it sounds, because captions become the conditioning signal during pretraining, and caption quality directly limits how controllable the final model is.

2. Train the tokenizer Compress
A tokenizer here is a separate neural network, an autoencoder, trained before anything else. An encoder squeezes a chunk of video into a small grid of numbers called a latent; a decoder reconstructs the video from it; training minimizes the difference. A typical tokenizer shrinks each spatial dimension about 8x and time 4 to 8x, so a 720p frame of roughly 2.7 million values becomes a grid of a few thousand latent positions, each a short list of numbers. For world models the tokenizer must be causal: the encoding of frame 10 can't depend on frame 11. The key trade-off is that more compression makes everything downstream cheaper, but detail the tokenizer discards (hand poses, small objects, text) can never be recovered. The tokenizer sets the ceiling on visual fidelity.

3. Pretrain on video Learn dynamics, part one
The large transformer trains on the curated latents, usually conditioned only on text captions or a starting image. Each training step follows the diffusion recipe: take a real future latent, add noise at a random level, and teach the model to recover the clean version given the noisy one plus the past. Without actions, the model learns "what typically happens next": appearance, object permanence, typical motion and gravity. This is where most of the compute goes, and it's the equivalent of LLM pretraining on internet text.

4. Action mid-training Learn dynamics, part two
The pretrained model is trained further on the action-labeled top of the pyramid, with each action injected as an extra input at every time step. This teaches what pretraining cannot: which future happens depends on what you do. Cosmos 3 frames the goal as learning the relationship in both directions: predicting the future from actions, inferring which actions explain an observed trajectory, and generating actions and future video together. Two details matter:

  • Action normalization. Every robot has different joints and ranges, so actions from different embodiments are converted into a common format and rescaled to a comparable range (roughly -1 to 1 in Cosmos 3).
  • Failure data. NVIDIA deliberately includes both successful and failed episodes. A world model that only ever saw successful grasps can't show a robot what a failed grasp looks like.

Where actions aren't logged, they are inferred, either with a latent action model (which learns its own compact action vocabulary from the differences between consecutive frames, as Genie does) or with an inverse dynamics model trained on a small labeled set and then used to pseudo-label a much larger unlabeled one.

5. Distill for real time Make fast
If pretraining used bidirectional attention, the model is converted to causal, frame-by-frame generation, and its 30 to 50 denoising steps are compressed to a few using the distillation and Self Forcing techniques covered earlier. This stage also covers serving engineering (caching, quantization, chunk sizes, hardware choice) that decides whether a session costs dollars or cents per hour.

6. Post-train. The general model is specialized: fine-tuned on a specific robot, a driving fleet, or a game's art style, and increasingly tuned with reward signals for physical correctness. An early form is reranking: generate several candidates and keep the one a learned critic scores best. Cosmos 3's reported physics results include a variant using best-of-N reranking with a learned world-model reward.

If every frontier lab runs roughly that pipeline, the obvious question is whether they're all building the same thing.

Is everyone doing the same thing? Where quality comes from

Convergence on the backbone

Broadly, yes. The field has converged quickly on a latent diffusion transformer, rolled out autoregressively, pretrained on massive unlabeled video, tuned on actions, and then distilled. This mirrors how language labs converged on the decoder-only transformer.

Where labs diverge

  • What the model outputs. Genie and Odyssey output video frames. World Labs bets on explicit 3D: its Atlas model was pretrained from scratch to natively operate on text, images, video, and 3D. AMI Labs outputs only embeddings and skips diffusion entirely.
  • The action space. Genie's actions are mostly navigation (move, look, trigger an event by prompt). Cosmos 3, Odyssey-3, and 1X target actual motor commands for robot bodies, a much harder and more physics-sensitive problem.
  • Starting point. DeepMind builds Genie on its Veo video model, inheriting years of video work. Others train from scratch to avoid inheriting a video model's habits (cinematic cuts, aesthetic bias) that hurt simulation.

Where quality differentiation comes from

Roughly in order of importance:

  1. Proprietary action-labeled data. The deepest moat, because it can't be bought or scraped. Waymo's fleet logs, Tesla's fleet, and robotics companies' deployment data are the top of the pyramid that internet video can't replace. Google, with YouTube-scale video and Waymo logs, is unusually well placed.
  2. Curation and captioning quality. Two labs with the same raw footage can produce very different models depending on filtering, deduplication, domain balance, and caption accuracy.
  3. Tokenizer quality. It caps fidelity: sharper hands, stable small objects, readable text.
  4. Long-horizon stability. Anti-drift training, memory design, and consistency tricks decide whether a world holds together for 10 seconds or 10 minutes. Short demos hide this.
  5. Compute and serving efficiency. World Labs reports Atlas's performance improves with training compute, so spending matters; at deployment, distillation and inference economics decide who can offer real-time worlds at a viable price.

Once language labs all adopted the transformer, the winners were decided by data, post-training, and infrastructure rather than core architecture. World models appear to be entering the same phase. The key difference: in language, the internet supplied nearly all the data anyone needed. In world models, the most valuable data, actions paired with consequences, mostly has to be generated by owning cars, robots, or games.

Part Three

Measuring & challenges

Part Two showed how world models are built. This part asks whether they work: how anyone checks that the physics is right, what's still unsolved, and where research is heading.

Measuring physics fidelity: the benchmark landscape

There's no single accepted benchmark. The useful way to understand the landscape is to sort benchmarks by how they judge physics, because each method catches some kinds of failure and misses others.

Reference-based: compare against real footage

Physics-IQ (Google DeepMind) is the most widely cited. It compares generated continuations against real recordings of 66 controlled physical experiments spanning solid mechanics, fluid dynamics, optics, thermodynamics, and magnetism, each filmed from three angles (198 test videos in all). Four metrics check whether motion happens in the right place, at the right time, and in the right amount, plus pixel-level error; the combined score is normalized so that two real recordings of the same experiment score 100%. It became the default number to report, including in the Sora 2 paper. On the original leaderboard, Cosmos 3 Super scores 59.7% when given several frames of context. Given the same single starting image as Sora 2, the comparison is much closer: 43.8% to Sora 2's 42.3%.

Blind spot: once everyone optimized for it, its flaws mattered. A June 2026 audit, Physics-IQ Verified, refined 57.6% of samples and improved 34.8% of prompts. On it, the same single-image comparison widens to 42.7% for Cosmos 3 Super's image-to-video version against 26.5% for Sora 2. More fundamentally, it rewards matching one real outcome, so a physically valid but different outcome is penalized.

Judge-based: a checklist of physics questions

  • WorldModelBench scores instruction following, common sense, and physics adherence across five violation types (Newton's laws, mass conservation, fluid dynamics, object penetration, gravity), over 350 instances in 7 domains.
  • PAI-Bench (NVIDIA, CVPR 2026) is broader: three tracks covering generation, conditional generation, and video understanding, with 2,808 real-world instances across driving, robotics, industry, and egocentric scenarios.

Blind spot: saturation. On WorldModelBench, Kling scores a perfect 1.00 on all five physics categories, which shows the test can't separate frontier models, not that physics is solved. These benchmarks also typically use vision-language models as judges, and those judges are themselves weak at spotting subtle physics errors.

Measurement-based: test the laws directly

World Models' Last Exam in Physics (October 2026) uses 40 controlled tasks spanning mechanics, optics, fluids, thermal and phase-change phenomena, electromagnetism, and surface tension, explicitly to test physical laws rather than similarity to a reference video.

Intuitive physics: does the model understand?

IntPhys 2 uses the "violation of expectation" method from infant psychology: show a possible and an impossible video and check whether the model registers surprise at the impossible one. This matters for non-generative models like V-JEPA, which can't be scored on pixels.

Closed-loop: does the world model make a robot better?

The category most relevant to robotics.

  • RBench scores task correctness and physical plausibility in robot–object interactions rather than visual realism, across 650 cases spanning dual-arm, humanoid, single-arm, and quadruped robots.
  • RoboWM-Bench, World-in-World, and WorldArena plug the world model into an actual control loop (closed-loop evaluation) and measure whether policies trained or planned inside it succeed.

The overarching gap

Nearly all these benchmarks were built for video generators, so they score passive continuations from an image or prompt. Very few test what makes a world model a world model: whether physics stays right in response to actions over long interactive rollouts. Closed-loop benchmarks come closest, and they're also the youngest.

Measurement is one unsolved problem among several. Here are the rest, many of them already met earlier in this post.

The biggest open challenges

  • Compounding error. Each predicted frame becomes the input for the next, so small mistakes snowball into drift (see exposure bias). This is the core reason interactive sessions are short; Genie 3 launched supporting only a few minutes of continuous interaction.
  • Memory and consistency. Keeping a world stable over long horizons without an ever-growing context. Genie 3's consistency was notable precisely because it emerged without being explicitly programmed.
  • Plausible vs. correct physics. The training objective rewards outputs that look like real video, not outputs that obey physics, and these diverge in rare or causal situations. A June 2026 benchmark study found current systems break down under small distribution shifts.
  • Action-labeled data scarcity. The top of the data pyramid.
  • Evaluation. No agreed score (see the benchmark landscape). Odyssey-3's preview, for example, initially shipped with no benchmark results; it has since self-reported a top score on Physics-IQ Verified, using its own prompts and best-of-8 selection.
  • Cost and latency. Real-time generation means serving a large model continuously per user.
  • Multi-agent interaction. DeepMind has acknowledged that modeling multiple independent agents in a shared environment remains difficult.

Each of those challenges has researchers working on it. These are the directions getting the most attention.

Frontier research directions

  • Explicit memory and 3D state. Rather than relying only on a history of pixels, recent work adds persistent spatial structure (e.g., "Beyond Pixel Histories: World Models with Persistent 3D State" and "MosaicMem," both 2026).
  • Unified world-action models. One network both imagines the future and outputs motor commands. Cosmos 3 can serve as a backbone for world-action models, and Odyssey-3 outputs world state plus motor actions. (For how world models differ from the vision-language-action models that decide what a robot does, see this earlier explainer.)
  • Predicting in representation space. The JEPA line, plus 2026 Dreamer variants that drop pixel reconstruction in favor of predicting representations.
  • Real-to-sim. Reconstructing real places as simulatable worlds. World Labs frames this as real-to-sim-to-real: a scalable engine for training and evaluating robot policies.
  • Closed-loop evaluation and multiplayer worlds. Benchmarks like World-in-World, and models like MIRA, a multiplayer interactive world model.
Part Four

Landscape

With the technology mapped, the last part turns to the market: where world models are used today, and who's building them.

Where world models are used today

  • Autonomous driving is the most mature deployment. In February 2026 Waymo announced the Waymo World Model, built on Genie 3, to simulate rare edge cases (a tornado, an elephant on the road) that can't be safely collected in reality.
  • Robotics: generating synthetic training data and evaluating policies before risking hardware. This is the core of NVIDIA's Cosmos pitch.
  • Creative tools and game prototyping: a visible consumer use, though still early.
  • Planning: the longer-term research bet, in which an agent imagines consequences before acting. This is the stated motivation of DeepMind and AMI Labs.

Those uses explain who's investing. These are the companies building the models themselves.

Who's building world models (as of October 2026)

Google DeepMind

Arguably the leader in interactive generative world models. Genie 3 (August 2025) was introduced as the first real-time interactive general-purpose world model. Project Genie opened it to the public in January 2026, limited to U.S. Google AI Ultra subscribers at $250/month, with generated worlds lasting 60 seconds. In May 2026, Genie gained the ability to simulate real streets using Street View. Genie 3 also underpins the Waymo World Model.

World Labs (Fei-Fei Li)

The spatial / 3D bet. Released Marble, its first commercial product, in late 2025, and a World API in January 2026. Raised $1 billion, acquired SceniX in July 2026 to push into robot training, and on September 1, 2026 launched Atlas, an omni world model pretrained from scratch on text, images, video, and 3D that stays 3D-consistent with everything it has seen.

Odyssey

Founded in late 2023 by Oliver Cameron and Jeff Hawke, both self-driving veterans. Raised a $310 million Series B at a $1.45 billion valuation, led by the venture arms of Amazon, NVIDIA, and AMD. Previewed Odyssey-3 on September 15, 2026, pitched as a single unified world model for robots, humanoids, cars, drones, and video games rather than one domain at a time. No independent benchmarks yet; its Physics-IQ results are self-reported.

AMI Labs (Yann LeCun)

The non-generative JEPA bet. Launched in March 2026 with a $1.03 billion seed round at a $3.5 billion pre-money valuation, the largest seed round in European history. Paris-headquartered and openly research-first, with a long time horizon. Released V-JEPA 2.1 alongside its launch.

NVIDIA

The open-weights infrastructure play. Cosmos 3 (June 2026) is an open "omnimodel" that understands and generates text, images, video, ambient sound, and actions, released in Super (64B), Nano (16B), and Edge (4B) sizes under an open license. Adopted across robotics (Doosan, LG, Samsung) and autonomous vehicles (Li Auto, Xiaomi), and backed by a Cosmos Coalition of partners including Runway and Skild AI.

Others worth tracking

  • Runway — GWM-1 (December 2025), extending its video stack toward interactive worlds.
  • Wayve — the GAIA family of driving world models.
  • 1X Technologies — a world model that serves as a virtual simulator for its humanoid robots.
  • Tesla — internal driving simulator built on fleet data.
  • Chinese labs — Tencent (HunyuanWorld) and Skywork (Matrix-Game), among others.

OpenAI

Shut down the Sora consumer app and API in March 2026. The Sora research team continues on world simulation research aimed at robotics and real-world physical tasks. No new world model has been publicly released from that effort.

Anthropic

No announced world model program. Anthropic's public focus has been language models, agents, and safety and interpretability research.

A closing caution. At AMI Labs' launch, its CEO predicted that within six months every company would call itself a world model to raise funding. That has largely come true. The quickest test for sorting real world models from rebranded video generators is the one from the start of this post: does it take actions, and does the world respond to them in real time?
Part Five

Reference

Terms used throughout, and where the facts come from.

Glossary

Action-conditioning
Generating the next state based on an input (action) that arrives during generation, rather than only on an initial prompt.
Autoregressive generation
Producing a sequence one piece at a time, each conditioned only on what came before.
Bidirectional attention
An attention pattern where every element can see every other, including future ones. Good for fixed clips; incompatible with interactivity.
Causal attention
An attention pattern where each element can see only the past. Required for real-time world models.
Closed-loop evaluation
Testing a world model by plugging it into a control loop and measuring whether an agent using it succeeds at real tasks.
Compounding error
Small prediction mistakes accumulating over a rollout because each output becomes the next input.
Diffusion
Generating a sample by starting from noise and refining it over many denoising steps.
Diffusion transformer
A transformer network trained to do diffusion's denoising; in world models it typically works on latents rather than raw pixels.
Distillation
Training a fast "student" model to reproduce the outputs of a slower "teacher" model, typically in far fewer steps.
Exposure bias
The mismatch between training on perfect past inputs and running on the model's own imperfect outputs; a primary cause of drift.
Inverse dynamics model (IDM)
A model that infers which action caused the change between two observations; used to pseudo-label unlabeled video.
JEPA (Joint Embedding Predictive Architecture)
Yann LeCun's approach of predicting abstract representations of missing or future content instead of reconstructing pixels.
KV cache
Stored intermediate computations for past context, so they don't need to be recomputed at every step.
Latent
A compact numerical representation of an image or video chunk produced by a tokenizer's encoder.
Latent action model
A model that learns its own compact vocabulary of actions from changes between consecutive frames, enabling training on video without action labels.
Rollout
A sequence of predictions produced by feeding a model's own output back in as its next input, step after step.
Tokenizer (video)
An autoencoder that compresses video into latents and reconstructs video from them; it caps the world model's visual fidelity.
World model
A learned system that predicts how the world's state changes in response to actions.

Sources