World Models, Video Generation, and Representation
LLM research is largely mature and has entered large-scale deployment and commercialization, while spatial intelligence remains an unsolved challenge for the AI community—and one that will take a long time to overcome.
Translated from Chinese with AI · Read the original
World Models, Video Generation, and Representation
Introduction
Most research on spatial intelligence falls into three areas:
- World Models: learning actions within and across video frames from video data, as in OpenAI’s early VPT work and Google DeepMind’s Genie series;
- Video Generation: synthesizing world states that never actually occurred—for example, the kinematic state of a cup after it is nudged—with Sora as a representative project;
- Representation: finding effective representations for compressing and modeling video and images without losing visual, motion, and other information, as in LeCun’s JEPA series;
At the data level, early research relied heavily on game data because game videos have several advantages:
- They are clean and consistent in maps, tasks, and colors, with relatively few OOD variations
- Data is easy to collect, unlike robotics companies that need many people to operate systems while wearing headsets
- Collection is efficient and can largely be completed online
- There is less dirty data and noise than in general internet data
Later, as applications generalized into areas such as robotics, research gradually shifted toward modeling with large volumes of internet video. Work in this direction has continued for roughly two years, yet many questions still lack consensus, including:
- No clear scaling law has emerged for world understanding. Here, this means zero-shot transfer across environments: can a model trained in game A perform comparably in game B? This is more about compositional generalization than simply memorizing trajectories in a model;
- Does video generation actually encode enough physical knowledge? MiniMax H3 and Seedance still produce many physically unrealistic results, even for common videos and prompts;
- How should multimodal data be represented? Pixel-only modeling is insufficient because, even when a representation can correctly encode and decode video, it often cannot expose real-world information in a form a robot can use;
The following sections review research on this topic from the past two years.
World Model
VPT
Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos, released by OpenAI in 2022 (OpenAI began researching video multimodality—and even robotics—much earlier than many people realize).
- Background: The internet contains vast amounts of gameplay and task videos, but they usually lack keyboard, mouse, or robot action labels and therefore cannot be used directly for BC (Behavior Cloning). Starting reinforcement learning from scratch also makes long-horizon skills extremely difficult to discover;
- Approach: Train an inverse dynamics model (IDM) on a small amount of action-labeled data, use it to add action labels to a much larger unlabeled video corpus, and then train a policy.
- Where do the labels come from? 1) Researchers hired people to play Minecraft while recording synchronized game video and keyboard actions from operation logs; 2) a model trained on that data then inferred actions between adjacent windows in internet videos to create pseudo-labels, which likely introduced some errors
- Action representation? A discrete-value action head is added directly to the Transformer;
- Model architecture: It consists of two models—an IDM for adding labels and a model for generating actions. The former can see the future, while the latter uses causal attention;
- Implementation: First, the IDM is trained on about 2,000 hours of Minecraft data containing keyboard and mouse records, learning to infer actions from frames before and after each action. It then generates pseudo-labels for roughly 70,000 hours of online video, which are used to train a behavior-cloning policy that depends only on historical observations. Finally, the model is fine-tuned on task data or with reinforcement learning.
- Results: The pretrained model can already chop down trees and craft planks and crafting tables. After reinforcement-learning fine-tuning, it produces a diamond pickaxe in 2.5% of ten-minute game episodes. Because this result includes online reinforcement learning, it should not be attributed to purely offline video learning.

Genie
Genie: Generative Interactive Environments
- Background: Traditional world models depend on paired observation-action data, while ordinary videos do not contain action records. Video-generation models also typically lack step-by-step control.
- Approach: Similar to VPT, discover latent actions between frames and then learn the visual outcomes produced by those actions;
- Tokenizer: The video ST-Transformer VQ-VAE encoder-decoder uses 1,024 codes, while actions use 8 codes;
- LAM (Latent Action Model): The encoder observes the frames before and after an action, learns a latent action, and compresses it into a codebook of 8 entries;
- Dynamics Model: It receives an action code and the previous frame, then predicts the next frame;
- Results: It can turn a 2D image into an interactive 2D environment.

My view: Compared with today’s end-to-end model training, this architecture seems somewhat dated. Compressing actions into a codebook of only 8 entries inherently loses information.
LAPA
Latent Action Pretraining From Videos
- Background: Released in 2024, this work addresses the high cost of robot teleoperation data, the lack of a unified action space across different robots, and the difficulty of directly reusing human video. In essence, it trains a video VQ-VAE encoder-decoder;
- Approach: It uses an unsupervised training method, similar to many existing WAM approaches
- Adjacent video frames are fed into a VQ-VAE to produce discrete action signals using a fixed-size codebook. After learning the action codebook, the action code and video frame are passed through a decoder to generate a video frame
- VQ-VAE: Vector Quantization, an encoder that maps continuous vectors to discrete signals
- The encoder and decoder here do not form a reconstruction model
- The trained VQ-VAE codebook and video frames are used for large-scale pretraining, yielding image+text -> action training data
- Fine-tune on a small amount of teleoperation data
- Adjacent video frames are fed into a VQ-VAE to produce discrete action signals using a fixed-size codebook. After learning the action codebook, the action code and video frame are passed through a decoder to generate a video frame

GR00T N1
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Background: Released by NVIDIA in 2025, GR00T N1 enables humanoid robots to understand vision and language, generate continuous actions, and adapt to different robot embodiments.
- Approach: Use a dual-system architecture that combines understanding and control, trained on a mixture of real robot data, human video, and synthetic data.
- Implementation: System 2 uses the Eagle-2 VLM to encode images and instructions; System 1 uses a DiT with flow matching to generate continuous action chunks. Each robot uses dedicated state and action adapter layers. Human video participates in training through LAPA-style latent actions, while synthetic data comes from video generation and simulated trajectory augmentation.
- This work also formally introduced the concept of the Data Pyramid for the first time

Isaac 0.5
Introducing Isaac 0.5: An Open-Source Embodied Foundation Model:
- A model released by Perception, with 36B parameters and an MoE architecture that dynamically activates parameters. Its defining feature is support for both video understanding and action generation
- Training
- Objectives: video understanding, spatial grounding, task-progress prediction, and action generation
- Predict task-relevant state changes in the scene (percepts)
Video Generation
Sora
Video generation models as world simulators
- OpenAI’s Sora, released in 2024, is still worth discussing. It launched with enormous expectations but lost momentum; in a few years, many people may barely remember it.
- Background: Long videos suffer from severe identity, motion, and spatial inconsistencies. Sora was the first long-video generation model.
- Approach: Images and video are tokenized—the tokenizer design was not disclosed—and then denoised with a large-scale Transformer-diffusion architecture. Training preserves videos at different resolutions and uses extensive text descriptions to improve instruction-following ability;
- The key idea was to transfer the DiT architecture used for images at the time to video and add spatiotemporal consistency, although the specific details were not disclosed
Genie 2
Genie 2: A large-scale foundation world model, a world model released by DeepMind in 2024, already resembles Marble, released by World Labs in 2026.
- Background: Agents need many environments for real interaction, so Genie 2 aims to provide Image-to-Environment generation, turning a 2D image into an interactive 3D environment.
- Approach: Similar to the methods above, it is trained with a causal Transformer. At every step, it takes historical frames and an action as input and generates the next frame;
- How does memory work? The report does not provide details. Given the short context length, the capability may simply emerge from the model’s pretraining data;
- Results: It demonstrates movement, opening doors, jumping, object interaction, and remembering scenes after they leave the field of view. A consistent environment can last up to about one minute, though most examples run for 10–20 seconds.
Dreamer 4
Training Agents Inside of Scalable World Models
- Background: Training an agent in a game requires a world model that can understand image+action -> next frame. The predicted next frame can then serve as a reward signal for agent RL—an arrangement that may be very similar to robotics.
- Approach:
- Tokenizer: The tokenizer is trained with image reconstruction and a perceptual loss, predicting randomly masked patches. It uses a causal Transformer encoder, and the encoding of each frame also uses information from the previous frame
- The network uses a Dynamics Transformer that receives historical visual representations, actions, noise, and other inputs, followed by block-causal attention. Tokens interact within each frame, while information cannot flow backward from the future along the time axis
- Shortcut Forcing: The denoising step size can also be predicted, and here
- Characteristics
- It can predict a wider range of interactive actions, such as chopping trees in Minecraft, thanks to finer-grained tokens and a longer context
Representation
RAE / VAE / V-RAE
- VAE (Auto-Encoding Variational Bayes): A very old concept that essentially learns the sampling distribution corresponding to the data;
Input image x ↓ EncoderOutput distribution parameters μ(x), σ(x) ↓ Sample from this distributionLatent variable z ↓ DecoderReconstructed image x̂- RAE (Diffusion Transformers with Representation Autoencoders): It uses pretrained visual representations to build an image latent space. In the era of large models, a pretrained model’s visual representation can be paired with a separately trained decoder to form a unified representation;
- V-RAE (V-RAE: Rethinking Video Latent Spaces for Generation): It applies the RAE idea to video, adding a temporal dimension and reconstruction metrics that preserve motion continuity.
DINO-WM
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
- Background: Pixel reconstruction can waste substantial model capacity on texture details. Task-specific rewards and policy training also limit the reusability of world models;
- Approach: This is broadly similar to Genie 2 and related systems, except that DINOv2 is used to encode images. Both approaches use latent or hidden states between adjacent frames as the representation for a dynamics model and learn two mappings:
- o(t) + z -> o(t+1): Predict the next state from the current state and an action
- o(t) + o(t+1) -> z: Predict the action from the change in state

V-JEPA 2 / V-JEPA 2.1
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Background: Pixel-by-pixel prediction must model many irrelevant and unpredictable details. A model needs to learn representations from observation that are useful for understanding and control.
- Approach: Predict masked video content in representation space, then add action conditioning with a small amount of robot-interaction data.
- Implementation: JEPA is pretrained on more than one million hours of internet video and images. A context encoder and predictor estimate the target encoder’s features. The visual representation is then frozen, and V-JEPA 2-AC is trained to predict future features from historical features and actions, using target images for MPC.

JiT
Back to Basics: Let Denoising Generative Models Denoise by Kaiming He
- Background: Current denoising processes obtain the final image by predicting noise, but noise is a high-dimensional and difficult prediction target. Predicting the target directly may turn the task into a lower-dimensional prediction problem, simplifying the model architecture while improving efficiency and accuracy.
- Approach:
- Manifold Assumption: A 256x256 image contains about 200,000 pixel values and can be viewed as high-dimensional data with roughly 200,000 dimensions, yet the actual semantic content—such as the outline of a cat—may occupy only a small part of that space
- Compared with a conventional DiT architecture, it makes two changes: 1) the input consists of patched pixels without any compression; 2) the output predicts the target rather than noise
