Intermediate level8 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain what a world model predicts, tell JEPA apart from models that generate video, and judge how far world models are from everyday robots.

The short answerA world model learns how an environment changes, so an agent can predict the results of actions and plan. Generative world models predict future frames; JEPA models predict abstract representations of hidden or future parts, and skip details that cannot be predicted.

In simple words

  • Ha and Schmidhuber (2018) trained an agent inside a learned model of a game; DreamerV3 later collected Minecraft diamonds from scratch.
  • Genie 3 generates worlds you can explore in real time, for a few minutes.
  • V-JEPA 2 has 1.2 billion parameters and learned from over 1 million hours of video; it does not generate video.
  • Robot results are early: two labs, 10 trials per task, about 16 seconds of planning per action.
Three things a model can learn to predictIllustration
  1. Next wordLanguage modelsPredict the next token of text.GPT, Claude, Gemini, Llama
  2. Next pictureGenerative world modelsPredict the next video frames or a whole 3D scene, in full visual detail.Genie 3, Cosmos, Marble
  3. Next meaningJEPAPredict a compact summary of a hidden or future part, and skip details that cannot be known.I-JEPA, V-JEPA 2
A simplified picture. JEPA stands for joint-embedding predictive architecture, which Yann LeCun proposed at Meta in 2022. The examples are named by their makers’ own descriptions.

Words to know

World model
An AI’s inner simulator that predicts what happens next, including after an action.
JEPA
Joint-embedding predictive architecture: it learns by predicting a summary of a hidden part, not its pixels or words.
Representation
The compact list of numbers a model uses to describe an input.

Learning by imagining

A world model predicts what happens next, including after an action. Ha and Schmidhuber (2018) trained an agent inside a compressed, learned model of a game, then moved it back to the real game.

DreamerV3 learns a model of its environment and improves by imagining futures. It was the first method to collect diamonds in Minecraft from scratch, without human data.

Generative world models

Google DeepMind’s Genie (2024) learned from unlabelled internet videos to make controllable 2D worlds. Genie 3 (2025) makes worlds from text at 24 frames a second, consistent for a few minutes, as a limited research preview.

Waymo says its World Model, built on Genie 3, generates camera and lidar scenes of rare events to test its self-driving software; it has shown examples, not measurements. NVIDIA’s Cosmos offers open-weight “world foundation models” for physical AI.

JEPA: predicting meaning

In 2022, Yann LeCun sketched how AI agents could plan with a world model that predicts what happens next. Its core, the joint-embedding predictive architecture or JEPA, turns two inputs, such as two moments of a video, into abstract representations, and predicts one from the other.

Meta argues that generative methods waste effort trying to fill in every pixel, though much is unpredictable. I-JEPA (2023) predicted representations of hidden image blocks, and V-JEPA (2024) learned from about 2 million videos by feature prediction alone.

From video to robot arms

V-JEPA 2 (2025) has 1.2 billion parameters and was trained on over 1 million hours of video. Adding under 62 hours of robot video let it steer robot arms in two labs without any data from them, guided by photos of the goal.

In 10 trials per task, it picked and placed a cup 80% of the time and a box 65%. Grasping a box with only one goal photo worked 25% of the time. It needed about 16 seconds of planning per action, and its authors list sensitivity to camera position and growing errors over long plans.

LeCun’s new company, AMI Labs, has reportedly raised $1.03 billion to build JEPA-based world models, with no product yet.

Try it yourself

Pause a video of a ball rolling towards a wall and say what happens next. You predict “it bounces back”, not every pixel: that is JEPA’s bet. Then open the knowledge graph and look at World Models.

Open the knowledge graph

Check yourself

  1. What is the main difference between Genie 3 and V-JEPA 2?

  2. Why does JEPA skip predicting every pixel?

  3. V-JEPA 2 moved robot arms in labs it had never seen. What should a buyer of robots conclude?

Back to the path

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.