Beginner level7 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain what a world model predicts, tell JEPA apart from models that generate video, and judge how far world models are from everyday robots.

The short answerA world model is an AI’s inner simulator: it predicts what will happen next, for example after a robot moves its arm. JEPA is one way to build it, which predicts the meaning of what comes next instead of every pixel.

In simple words

  • Chatbots predict the next word; world models predict how a scene or a situation will change.
  • Some world models, such as Google DeepMind’s Genie 3, generate whole worlds you can move around in.
  • JEPA, an idea from Yann LeCun at Meta, predicts a short summary instead of drawing pictures.
  • Robots that use world models still work mainly in labs, in small tests.
Three things a model can learn to predictIllustration
  1. Next wordLanguage modelsPredict the next token of text.GPT, Claude, Gemini, Llama
  2. Next pictureGenerative world modelsPredict the next video frames or a whole 3D scene, in full visual detail.Genie 3, Cosmos, Marble
  3. Next meaningJEPAPredict a compact summary of a hidden or future part, and skip details that cannot be known.I-JEPA, V-JEPA 2
A simplified picture. JEPA stands for joint-embedding predictive architecture, which Yann LeCun proposed at Meta in 2022. The examples are named by their makers’ own descriptions.

Words to know

World model
An AI’s inner simulator that predicts what happens next, including after an action.
JEPA
Joint-embedding predictive architecture: it learns by predicting a summary of a hidden part, not its pixels or words.
Representation
The compact list of numbers a model uses to describe an input.

An inner simulator

Before you push a glass towards the edge of a table, you can picture it falling. A world model gives an AI that kind of foresight. It learns how things change, so it can try out actions in its head before it acts.

The idea is decades old. In 2018, for example, researchers trained an AI to play a game inside its own “dream” of the game, then moved it back to the real one.

Two ways to predict

One way is to generate the future as pictures. Google DeepMind’s Genie 3 turns a text description into a world you can move around in, at 24 frames a second, for a few minutes. Such worlds can still break the rules of real physics.

The other way is JEPA, which Yann LeCun proposed at Meta in 2022. Think of covering half of a photo: you guess “a table, maybe a chair”, not every pixel. JEPA learns the same way, predicting a short summary of the hidden part, called a representation, and skipping details nobody could know.

Where they are today

Meta’s V-JEPA 2, released in 2025, learned from over a million hours of video. With a little robot video added, it steered robot arms in two labs it had never seen. In tests of picking up an object and putting it down, it succeeded 65 to 80% of the time, over 10 tries per task.

That is real progress, but small tests in labs are not robots in homes. Each move took about 16 seconds of planning. Meta also found that every AI model it tested, even the best, V-JEPA 2, did little better than guessing on its new physics test, while people were nearly perfect.

Try it yourself

Pause a video of a ball rolling towards a wall and say what happens next. You predict “it bounces back”, not every pixel: that is JEPA’s bet. Then open the knowledge graph and look at World Models.

Open the knowledge graph

Check yourself

  1. What does a world model try to predict?

  2. You cover half of a photo and guess what is behind it. Which approach matches how you guess?

  3. Meta’s V-JEPA 2 moved robot arms in two labs. What does that show?

Back to the path

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.