Intermediate level8 min readTo learn at another level, choose it before you start the course.
After this lesson you canExplain what a world model predicts, tell JEPA apart from models that generate video, and judge how far world models are from everyday robots.
The short answerA world model is an AI’s inner simulator: it predicts what will happen next, for example after a robot moves its arm. JEPA is one way to build it, which predicts the meaning of what comes next instead of every pixel.
In simple words
- Chatbots predict the next word; world models predict how a scene or a situation will change.
- Some world models, such as Google DeepMind’s Genie 3, generate whole worlds you can move around in.
- JEPA, an idea from Yann LeCun at Meta, predicts a short summary instead of drawing pictures.
- Robots that use world models still work mainly in labs, in small tests.
The short answerA world model learns how an environment changes, so an agent can predict the results of actions and plan. Generative world models predict future frames; JEPA models predict abstract representations of hidden or future parts, and skip details that cannot be predicted.
In simple words
- Ha and Schmidhuber (2018) trained an agent inside a learned model of a game; DreamerV3 later collected Minecraft diamonds from scratch.
- Genie 3 generates worlds you can explore in real time, for a few minutes.
- V-JEPA 2 has 1.2 billion parameters and learned from over 1 million hours of video; it does not generate video.
- Robot results are early: two labs, 10 trials per task, about 16 seconds of planning per action.
The short answerWorld models learn environment dynamics for prediction and planning: generative ones model observations such as frames or 3D scenes, while JEPA models predict in representation space and discard unpredictable detail. The evidence is strongest in games and simulation; robot results remain small, slow and fragile.
Key points
- Ha and Schmidhuber (2018) and DreamerV3 (2023; Nature, 2025): learning and acting inside learned models.
- LeCun (2022) sketched agents built on stacked JEPA modules that learn without labels; it is a proposal, not a system.
- V-JEPA 2-AC: under 62 hours of robot video, then zero-shot planning on robot arms, 80% and 65% at pick-and-place.
- On IntPhys 2, every tested model scored below 60%, near chance (50%); the best, V-JEPA 2, reached 57.5%, and people about 96%.
- Next wordLanguage modelsPredict the next token of text.GPT, Claude, Gemini, Llama
- Next pictureGenerative world modelsPredict the next video frames or a whole 3D scene, in full visual detail.Genie 3, Cosmos, Marble
- Next meaningJEPAPredict a compact summary of a hidden or future part, and skip details that cannot be known.I-JEPA, V-JEPA 2
Words to know
- World model
- An AI’s inner simulator that predicts what happens next, including after an action.
- JEPA
- Joint-embedding predictive architecture: it learns by predicting a summary of a hidden part, not its pixels or words.
- Representation
- The compact list of numbers a model uses to describe an input.
An inner simulator
Before you push a glass towards the edge of a table, you can picture it falling. A world model gives an AI that kind of foresight. It learns how things change, so it can try out actions in its head before it acts.
The idea is decades old. In 2018, for example, researchers trained an AI to play a game inside its own “dream” of the game, then moved it back to the real one.
Two ways to predict
One way is to generate the future as pictures. Google DeepMind’s Genie 3 turns a text description into a world you can move around in, at 24 frames a second, for a few minutes. Such worlds can still break the rules of real physics.
The other way is JEPA, which Yann LeCun proposed at Meta in 2022. Think of covering half of a photo: you guess “a table, maybe a chair”, not every pixel. JEPA learns the same way, predicting a short summary of the hidden part, called a representation, and skipping details nobody could know.
Where they are today
Meta’s V-JEPA 2, released in 2025, learned from over a million hours of video. With a little robot video added, it steered robot arms in two labs it had never seen. In tests of picking up an object and putting it down, it succeeded 65 to 80% of the time, over 10 tries per task.
That is real progress, but small tests in labs are not robots in homes. Each move took about 16 seconds of planning. Meta also found that every AI model it tested, even the best, V-JEPA 2, did little better than guessing on its new physics test, while people were nearly perfect.
Try it yourself
Pause a video of a ball rolling towards a wall and say what happens next. You predict “it bounces back”, not every pixel: that is JEPA’s bet. Then open the knowledge graph and look at World Models.
Open the knowledge graphCheck yourself
What does a world model try to predict?
You cover half of a photo and guess what is behind it. Which approach matches how you guess?
Meta’s V-JEPA 2 moved robot arms in two labs. What does that show?
Words to know
- World model
- An AI’s inner simulator that predicts what happens next, including after an action.
- JEPA
- Joint-embedding predictive architecture: it learns by predicting a summary of a hidden part, not its pixels or words.
- Representation
- The compact list of numbers a model uses to describe an input.
Learning by imagining
A world model predicts what happens next, including after an action. Ha and Schmidhuber (2018) trained an agent inside a compressed, learned model of a game, then moved it back to the real game.
DreamerV3 learns a model of its environment and improves by imagining futures. It was the first method to collect diamonds in Minecraft from scratch, without human data.
Generative world models
Google DeepMind’s Genie (2024) learned from unlabelled internet videos to make controllable 2D worlds. Genie 3 (2025) makes worlds from text at 24 frames a second, consistent for a few minutes, as a limited research preview.
Waymo says its World Model, built on Genie 3, generates camera and lidar scenes of rare events to test its self-driving software; it has shown examples, not measurements. NVIDIA’s Cosmos offers open-weight “world foundation models” for physical AI.
JEPA: predicting meaning
In 2022, Yann LeCun sketched how AI agents could plan with a world model that predicts what happens next. Its core, the joint-embedding predictive architecture or JEPA, turns two inputs, such as two moments of a video, into abstract representations, and predicts one from the other.
Meta argues that generative methods waste effort trying to fill in every pixel, though much is unpredictable. I-JEPA (2023) predicted representations of hidden image blocks, and V-JEPA (2024) learned from about 2 million videos by feature prediction alone.
From video to robot arms
V-JEPA 2 (2025) has 1.2 billion parameters and was trained on over 1 million hours of video. Adding under 62 hours of robot video let it steer robot arms in two labs without any data from them, guided by photos of the goal.
In 10 trials per task, it picked and placed a cup 80% of the time and a box 65%. Grasping a box with only one goal photo worked 25% of the time. It needed about 16 seconds of planning per action, and its authors list sensitivity to camera position and growing errors over long plans.
LeCun’s new company, AMI Labs, has reportedly raised $1.03 billion to build JEPA-based world models, with no product yet.
Try it yourself
Pause a video of a ball rolling towards a wall and say what happens next. You predict “it bounces back”, not every pixel: that is JEPA’s bet. Then open the knowledge graph and look at World Models.
Open the knowledge graphCheck yourself
What is the main difference between Genie 3 and V-JEPA 2?
Why does JEPA skip predicting every pixel?
V-JEPA 2 moved robot arms in labs it had never seen. What should a buyer of robots conclude?
Model-based learning
Ha and Schmidhuber (2018) learned a compressed model of a game and its dynamics, trained a controller inside that “dream”, and transferred it back. DreamerV3 (Hafner and others, 2023; Nature, 2025) learns a world model and improves by imagining futures.
DreamerV3 was the first method to collect Minecraft diamonds from scratch, without human data.
The JEPA proposal
LeCun’s 2022 position paper sketches how an autonomous agent could learn and plan. In it, stacked JEPA modules learn without labels, a world model lets the agent try plans before acting, and built-in drives set its goals. It is a proposal, not a finished system.
In a JEPA, encoders map two inputs to representations, and a predictor estimates one from the other, so hard-to-predict detail can be dropped. I-JEPA (Assran and others, 2023) predicted hidden image-block representations from a visible block without hand-made augmentations, training a ViT-Huge on 16 GPUs in under 72 hours.
V-JEPA (Bardes and others, 2024) learned from about 2 million videos by feature prediction alone, with no text and no pixel reconstruction, reaching 81.9% on Kinetics-400. Meta claims 1.5 to 6 times better training efficiency than generative methods.
V-JEPA 2 and planning
V-JEPA 2 (2025) has 1.2 billion parameters, trained on over 1 million hours of video and 1 million images. V-JEPA 2-AC added under 62 hours of unlabelled DROID robot video and planned zero-shot on Franka arms in two labs, steered by goal images.
Over 10 trials per task per lab, pick-and-place succeeded 80% of the time with a cup and 65% with a box, against 15% and 10% for the Octo baseline. Planning took about 16 seconds per action. The authors cite sensitivity to camera position, compounding errors over long horizons and the need for goal images.
Meta notes that V-JEPA 2 predicts at one time scale only. On Meta’s IntPhys 2 benchmark, every tested model scored below 60%, near chance (50%): the best, V-JEPA 2, reached 57.5%, and people about 96%. V-JEPA 2.1 (2026) claims a 20-point gain in real-robot grasping.
Generative world models
Genie (Bruce and others, 2024), with 11 billion parameters, learned latent actions from unlabelled internet video. Genie 3 (2025) renders text-prompted worlds at 720p and 24 frames a second, consistent for a few minutes; its listed limits include few agent actions and inexact real places.
Project Genie opened Genie 3 to US Google AI Ultra subscribers in January 2026, with 60-second worlds that may not follow real physics. Waymo’s World Model builds rare driving scenes on Genie 3, shown by example without metrics.
NVIDIA’s Cosmos (2025) offers open-weight world foundation models with tokenizers and data tools, and World Labs’ Marble builds 3D worlds that can be exported. As our news reported, AMD said it will buy World Labs.
Try it yourself
Pause a video of a ball rolling towards a wall and say what happens next. You predict “it bounces back”, not every pixel: that is JEPA’s bet. Then open the knowledge graph and look at World Models.
Open the knowledge graphCheck yourself
In a JEPA, where is the prediction loss computed?
Which limit did V-JEPA 2-AC’s authors report for robot planning?
On IntPhys 2, every tested model, V-JEPA 2 included, scored close to chance, while people were nearly perfect. What does this suggest?
Sources
- World Models (Ha and Schmidhuber, 2018, arXiv)
- Mastering Diverse Domains through World Models, DreamerV3 (Hafner and others, 2023, arXiv)
- Mastering diverse control tasks through world models (Hafner and others, 2025, Nature)
- A Path Towards Autonomous Machine Intelligence (LeCun, 2022, OpenReview)
- Yann LeCun on a path towards human-level AI (Meta, February 2022)
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture, I-JEPA (Assran and others, 2023, arXiv)
- I-JEPA: the first AI model based on LeCun’s vision (Meta, June 2023)
- Revisiting Feature Prediction for Learning Visual Representations from Video, V-JEPA (Bardes and others, 2024, arXiv)
- V-JEPA: video learning without generating pixels (Meta, February 2024)
- V-JEPA 2 world model and physics benchmarks (Meta, June 2025)
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (Assran and others, 2025, arXiv)
- V-JEPA 2.1 (2026, arXiv)
- IntPhys 2: benchmarking intuitive physics in video models (Meta, 2025, arXiv)
- Yann LeCun’s AMI Labs raises $1.03 billion (TechCrunch, March 2026)
- Genie: Generative Interactive Environments (Bruce and others, 2024, arXiv)
- Genie 3: a new frontier for world models (Google DeepMind, August 2025)
- Project Genie (Google, January 2026)
- The Waymo World Model (Waymo, February 2026)
- Cosmos World Foundation Model Platform for Physical AI (NVIDIA, 2025, arXiv)
- Marble (World Labs, November 2025)
- AMD will buy Fei-Fei Li’s World Labs (Silicon AI News)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.