Why AI Is Being Taught to Simulate the World

Odyssey has introduced Odyssey-2 Max, its largest world model. Formally, it is a relative of video generators, but the logic is different: the system does not assemble a ready-made clip from a prompt, but step by step predicts the next state of the scene and allows you to continue the simulation in real time.

The company is not selling "beautiful video," but a controllable environment where the future of a frame depends on the previous state and the user's action. For games, this sounds like an interactive scene. For robotics, like a training environment. For defense and medicine, like a reason to carefully examine the quality of verification, because beautiful physics in a demo does not yet equal a reliable model of cause-and-effect relationships.

The release includes several specific numbers. 2 Max is roughly 3 times larger than 2 Pro and was trained with a 10-fold increase in computation. On the physical section of VBench 2, the result increased from 49.67 to 58.52. On PAI-Bench Physics, from 91.67 to 93.02. The company also claims that all shown simulations ran in real time and could continue for more than 120 seconds.

Technically, the model is built as an autoregressive diffusion transformer. Important details: long context, causal attention, action control through latent representations, flow matching in continuous latent space, and reduction of denoising steps. Training was done on several hundred NVIDIA B200s.

The most interesting part here is not the picture. The creators draw an analogy with language models: predicting the next token gave systems the ability to imitate reasoning, and predicting the next state of the world should give physical intuition. This is an ambitious hypothesis, and it explains well why there is so much investor and research interest in world models right now.

But all this needs to be verified carefully. VBench and PAI-Bench evaluate the consistency of generated video, not the system's suitability for real robotics or scientific modeling. Stable background, smooth motion, and plausible mechanics are useful, but they do not prove that the model understands causal relationships in a strict sense.

The comparison with Sora, Veo, Kling, and Runway is also arranged favorably for the developers of 2 Max. These systems are excluded from the table as bidirectional video models because they are not designed for interactive prediction of future states. The argument is logical, but the comparison field becomes smaller: we are talking about a category that the company itself is trying to establish as separate.

Another point: the model is available in a private beta for partners. This means independent verification is still limited. The main questions will arise in long scenarios where the user makes strange actions, the scene gradually accumulates errors, and physical plausibility begins to conflict with controllability.

The release is still significant. Generative video is gradually splitting into two lines: production of ready-made visual content and interactive environment simulators. The first line serves media. The second could become the basis for simulators, games, agents, robotics, and planning systems.

The race for world models has become a separate direction. It competes not so much on the beauty of frames, but on the stability of causality, simulation horizon, controllability, and the cost of real-time generation.


❗️❗️❗️❗️❗️❗️❗️❗️ / Not banned in the Russian Federation