Veo 3 generated buoyancy simulations and maze-solving sequences across 18,000+ videos without being trained on either task. This is emergent physics: physical-reasoning behaviour that surfaces in generated output without explicit training. No top-15 text-to-video model surpasses 67% on the formal benchmark. Both observations are real, and they describe different things.

Chart: VBench-2.0 total score, top text-to-video models, early 2026.

What VBench-2.0 measures: 18 dimensions across motion dynamics, semantic alignment, temporal consistency, and physical plausibility. The benchmark scores text-to-video model outputs against human-rated reference distributions for each dimension. A 67% aggregate score means the model's output distribution matches the reference distribution on roughly two-thirds of the rated dimensions. Veo 3 leads at the top of the leaderboard. The next-tier models (Wan, Sora, Kling) sit within a 10-point band below.

What the emergent-physics observation describes: Veo 3 produced videos demonstrating buoyancy (a ball thrown into water sinks, displaces water, bubbles rise) and maze-solving (a generated character navigates through a generated maze without getting stuck on walls) across 18,000+ examples. These were not training tasks. The model was trained on text-to-video generation broadly. The behaviour appeared.

Two interpretations of the gap between these observations.

The first: emergent physics is real but narrow. The model has learned regularities about how objects move, how light behaves, how characters navigate space. These regularities appear in some generated outputs. They do not constitute a general model of physics: the model fails on physics scenarios it has not seen analogues of in training, and the failures are visible in the formal benchmark scores. 33% of rated dimensions are not matching reference distributions.

The second: the formal benchmark and the emergent-physics observation are measuring different things, and both are valid. VBench-2.0 measures aggregate match to reference distributions across many dimensions, with no special weighting for the dimensions that represent physical reasoning. A model can score 67% overall while being exceptionally good on the subset of dimensions that capture emergent physics. The aggregate score under-reports the specific capability that the qualitative observation highlights.

Both interpretations are supported by the available data. The benchmark scores show the model is not at human reference distributions across all dimensions. The 18,000-example observation shows the model demonstrates physical-reasoning behaviour at a scale that is not consistent with random chance. These observations co-exist.

The data observation: text-to-video generation models in early 2026 demonstrate physical-reasoning behaviour that exceeds what their architecture was designed to produce, while remaining below the aggregate quality threshold that formal benchmarks measure. The two patterns are not contradictory: they describe different aspects of model behaviour.

For practitioners evaluating text-to-video generation capability for enterprise use cases, this surfaces a specific evaluation question: which aspect of capability matters for your workload? If the workload requires aggregate quality matching reference distributions across many dimensions (advertising, brand-aligned content production), VBench-2.0 and similar benchmarks answer the relevant question. If the workload requires specific physical-reasoning capabilities (simulation, training data generation for embodied agents, scenario generation for testing), the benchmark scores are insufficient and capability needs to be evaluated against the specific physical-reasoning task.

Three observations follow from the data.

The first: the rate of capability gain on text-to-video benchmarks has been faster than expected. Veo 3, Sora 2, and Wan 2.5 represent a year-on-year capability shift on the order of multiple percentage points across the aggregate score, with disproportionate gains on the dimensions related to physical plausibility. The trajectory has not plateaued.

The second: the emergent-physics behaviour suggests text-to-video models are picking up regularities from training data that go beyond pixel prediction. What "physical reasoning" actually consists of computationally (whether it is genuine model-internal representation of physical laws, or pattern matching against video sequences that happen to demonstrate those laws) is not resolved by the available data. Both interpretations remain plausible.

The third: the gap between aggregate benchmark scores and qualitative observations of specific capabilities is now a recurring pattern across multiple modalities. The same pattern shows up in mathematical reasoning (high benchmark scores on some maths; jagged failures on adjacent maths) and in agent capability (strong performance on some tasks; conspicuous failures on others). The pattern is consistent enough that it should now be expected in capability evaluations, not treated as anomalous when it appears.

The 67% aggregate score and the 18,000-example emergent-physics observation are both real. They are evidence of capability at different levels of resolution. The procurement and deployment questions that follow are workload-specific. The data is in. Whether your frameworks use it to interrogate text-to-video capability at the right level of resolution is the choice in front of you.

Sources

  • Primary: Stanford AI Index 2026, Chapter 2 (Technical Performance) 2.3 — hai.stanford.edu/ai-index/2026
  • Video benchmark data: VBench-2.0 Leaderboard — multimodal video evaluation
  • Emergent-physics observations: Google DeepMind Veo 3 technical report; cross-vendor qualitative analysis across Sora, Wan, Kling