CaliBench
CaliBench is a benchmark for image-to-video generative models that asks whether a model reproduces the correct distribution of physical outcomes across many generations from the same starting frame, rather than whether any single generated video looks physically plausible. It was built by Odyssey, the world model company founded by Oliver Cameron and Jeff Hawke, and announced on the company blog on August 10, 2026 in a post bylined Jon Sadeghi [1]. The accompanying paper, "CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?", was posted to arXiv on 17 August 2026 (revised 5 September 2026) and was accepted at Transactions on Machine Learning Research [2]. Its authors are Jonathan Sadeghi, Jenny Seidenschwarz, Jesse Allardice, Sirish Srinivasan, Benjamin Graham and Jeffrey Hawke, all at Odyssey [2].
The benchmark scores nine physical scenes whose outcome distributions are known exactly in closed form, such as a Galton board (binomial over landing bins) or a fair die (uniform over six faces). A model generates 32 videos per scene from one fixed conditioning frame, varying only the seed; a vision-language model reads the discrete outcome of each video; and the empirical distribution is compared with the analytic reference. The headline finding is that all six frontier systems tested pile probability mass onto a handful of outcomes instead of reproducing the reference, with one model producing the same die face on every seed [2].
Calibration versus per-sample realism
The argument behind CaliBench is that a video model standing in for physical reality has two distinct jobs, and existing evaluations only test one of them. Most video benchmarks either judge each individual generation (sharpness, temporal consistency, frame-by-frame physical plausibility) or compare whole datasets in a learned feature space, as Frechet Video Distance does [1][4]. Neither answers the question of whether many generations from the same starting state recover the right spread of outcomes.
The paper's framing is that macroscopic physics is effectively deterministic, and what reads as randomness in a dice roll or a roulette spin is sensitive dependence on initial conditions that a single low-resolution conditioning frame cannot resolve. A model that maps that one macro-state to a single trajectory every time is not simulating the system; it is reciting its most likely output. A model can therefore pass every per-sample check and still be badly miscalibrated [2].
Odyssey frames the stakes in terms of stress-testing autonomous systems. If a world model is used to exercise a self-driving stack or a robot policy in simulation, a simulator that collapses onto the single likeliest future hides exactly the rare tail events the simulation was built to surface, and returns a world that is more deterministic, and safer, than the real one [1][2]. The same concern appears in the paper's discussion of autonomous driving and robotics as safety-critical deployment settings for physical AI [2].
Benchmark design
CaliBench's nine scenes fall into four groups by the structure of their reference distribution. Every scene pairs a single conditioning image with a fixed text prompt, and the reference is derived analytically rather than estimated from data [2][3].
| Scene | Outcome read from the video | Outcomes (k) | Reference distribution |
|---|---|---|---|
| Physical Galton board | Landing bin index | 11 | Binomial(10, 1/2), photorealistic board with a metallic ball |
| Animated Galton board | Landing bin index | 11 | Binomial(10, 1/2), 2D vector illustration with a green ball |
| Ball fork | Left or right channel of a symmetric Y-ramp | 2 | Bernoulli(1/2) |
| Walking | Left or right turn at a symmetric T-junction | 2 | Bernoulli(1/2) |
| Double pendulum | Side of the vertical axis the outer mass rests on | 2 | Bernoulli(1/2) |
| Dice | Upward face of a six-sided die | 6 | Uniform over 6 |
| Cards | Suit of the drawn card | 4 | Uniform over 4 |
| Lottery | Number on the ejected ball (1-20) | 20 | Uniform over 20 |
| Roulette | Pocket colour on a European wheel | 3 | 18/37 red, 18/37 black, 1/37 green |
The two Galton boards share an identical analytic target but differ in visual domain, which tests whether calibration survives a change of rendering style. Both use a 10-row lattice, so the central bin carries a probability of about 24.6 percent, the number several models overshoot [2].
Two scenes rely on chaotic dynamics for their reference to hold. The roulette prompt asks for a rapidly spinning wheel so that the ball deflects repeatedly off the frets, and the double pendulum is released from a high-energy configuration. The authors did not simply assume the pendulum's Bernoulli(1/2) marginal: they integrated the nonlinear double-pendulum equations from the conditioning frame's measured release state (roughly 146 and 215 degrees for the two angles, length ratio 0.82, equal masses, released from rest) over an ensemble of small perturbations, and found the simulated probability of the outer mass resting left stays within 0.05 of 0.50 across the 3 to 5 second range of model video lengths [2]. The check is shipped in the released code as verify_pendulum_v3.py [3].
Generation and extraction protocol
For each scene and model pair, CaliBench produces 32 videos from identical inputs, changing only the integer seed drawn from a fixed set. Evaluating one model across all nine scenes therefore costs 288 generations, and the full six-model matrix in the paper required 1,728 video rollouts and 5,184 vision-language-model queries [2]. All other settings are held constant per model at the provider's defaults, including classifier-free guidance scale and prompt expansion. Durations differ between providers because each model was run at or near the shortest length its API offers (WAN-2.7 and SeeDance-2.0, which accept any integer duration, were set slightly above their minimums), and output resolutions were left at each provider's setting [2].
| Model | API identifier | Duration | Resolution |
|---|---|---|---|
| WAN-2.7 | wan-video/wan-2.7-i2v | 3 s | 720p |
| SeeDance-2.0 | bytedance/seedance-2.0 | 4 s | 480p |
| HappyHorse-1.0 | alibaba/happyhorse-1.0 | 3 s | 720p |
| Veo 3.1 | google/veo-3.1 | 4 s | 720p |
| Runway Gen-4.5 | runwayml/gen-4.5 | 5 s | not exposed by the API |
| Cosmos3-Super | nvidia/Cosmos3-Super (run locally) | 5.04 s (121 frames at 24 fps) | 1280x704 |
Outcomes are extracted automatically with Gemini 3.1 Pro [14], prompted per scene to return a bin index, a die face, a suit, a colour, or a null token when the final state is unreadable. Each query is run three times and aggregated by majority vote, with ties resolved to null [2]. The null rules are explicit: a roulette ball balanced on a fret is null, a lottery tube containing anything other than exactly one ball is null, a die must show a single identifiable face, and a pendulum whose rods deform or fracture is null.
The extraction pipeline was validated against a human annotator on a stratified sample of 18 videos per scene, 162 in total. Agreement was 152 of 162, about 93.8 percent, with a false-null rate of 4 of 107 human-labelled valid outcomes (3.7 percent) and a missed-null rate of 4 of 55 human-labelled unscorable sequences (7.3 percent). The physical Galton board and dice scenes reached perfect agreement; cards was worst at 83 percent [2]. The blog post rounds this to agreement of "around 93%" [1].
Metrics
CaliBench deliberately separates two failures that a single accuracy score would blur together [1][2]:
- Scorability (rho) is the fraction of the 32 generations that yield a clean readable outcome at all. A ball that settles in one bin is scorable; a video that spawns three balls, or a pendulum that tears itself apart, is not.
- Calibration is the total variation distance between the empirical outcome distribution over the scorable generations and the analytic reference, read as the fraction of probability mass misallocated. Total variation was chosen over a Wasserstein distance because several outcome spaces (card suits, roulette colours) are unordered categories with no natural inter-outcome metric.
Significance is assessed with a Monte Carlo exact chi-squared goodness-of-fit test using 50,000 null replicates per cell, rather than the asymptotic chi-squared reference, because expected counts fall below Cochran's rule on four scenes (the Galton board tails, the rare green roulette pocket, and every bin of the 20-way lottery) [2]. Because calibration is the null hypothesis, the test can only ever evidence miscalibration, never confirm calibration, and at 32 generations per cell it reliably detects only large deviations. Across the family of testable cells the authors control the false discovery rate with the Benjamini-Hochberg procedure at alpha = 0.05 [2].
The summary metric is mean normalised total variation (mnTV). Raw total variation is biased upward at finite sample sizes and the bias grows with the number of outcomes, so raw values are comparable within a scene but not across scenes. mnTV subtracts each cell's closed-form null floor, folds in unscorability as a penalty, and averages the result over all nine scenes. It is bounded above by 1 and approaches 0 only when a model is both fully scorable and calibrated up to sampling noise; a perfectly calibrated synthetic operator lands in a central 95 percent range of roughly -0.04 to 0.04 [2]. The paper is explicit that mnTV is meant to sit beside the per-scene grid rather than replace it, since averaging over heterogeneous scenes masks per-scene reversals and conflates the two axes the rest of the work treats as orthogonal.
Results
The paper evaluated six image-to-video systems. Aggregate mnTV scores, with bootstrap 95 percent intervals from 50,000 resamples, were [2]:
| Model | Developer | mnTV (lower is better) | Bootstrap interval |
|---|---|---|---|
| SeeDance-2.0 | ByteDance | 0.39 | 0.38 to 0.48 |
| Runway Gen-4.5 | Runway | 0.48 | 0.48 to 0.57 |
| HappyHorse-1.0 | Alibaba | 0.54 | 0.51 to 0.59 |
| WAN-2.7 | Alibaba (Wan team) | 0.58 | 0.54 to 0.63 |
| Cosmos3-Super | NVIDIA | 0.63 | 0.62 to 0.70 |
| Veo 3.1 | Google DeepMind | 0.65 | 0.63 to 0.69 |
Every score sits far above the roughly zero that a calibrated model would reach. SeeDance-2.0 has the lowest penalty, but the paper reads the overlapping bootstrap intervals as broad performance tiers rather than a strict ranking [2]. The blog post puts the same point more loosely, saying that Seedance performs best overall while HappyHorse is better on some individual scenes [1].
The main structural result is heterogeneity. Across the eight informative scenes no model dominates: Runway Gen-4.5 leads on the animated Galton board, SeeDance-2.0 on the physical board and the lottery, HappyHorse on the ball fork, Cosmos3-Super on walking and dice, and Veo 3.1 on roulette, while cards is a three-way tie on total variation [2]. The authors report finding no simple property of a model or a scene that predicts calibration performance [1].
Specific failure patterns:
- Mode collapse is the dominant failure. Of the 54 scene-model cells, 44 had enough valid generations to compute a chi-squared p-value, and 28 of those remained significantly miscalibrated after Benjamini-Hochberg correction. Every cell with severe collapse (raw p below 0.001) survives even the stricter Bonferroni bound [2].
- Veo 3.1 collapsed completely on dice, returning the same face across every seed [1][2].
- The Galton boards draw too much mass onto the centre. Most models over-concentrate on the central bin well past its true 24.6 percent peak, although on the animated board SeeDance-2.0 fails in the opposite direction, overloading the left tail [2].
- The double pendulum broke every model. No system produced more than three scorable generations out of 32 (Cosmos3-Super three, SeeDance-2.0 two, the rest at most one), which the paper calls a universal structural failure and excludes from per-scene distributional claims [2].
- Roulette is often unreadable. Generated videos frequently leave the ball ambiguously placed, so several models score low on scorability there rather than on calibration [2].
- Calibration does not transfer across distribution families. Models close to calibrated on the binary symmetric scenes degrade sharply on high-cardinality targets such as the 20-way lottery or the skewed roulette reference [2].
Two ablations rule out the obvious confounders. Regenerating SeeDance-2.0 at 720p instead of its benchmarked 480p leaves the strongly miscalibrated scenes unchanged, and re-running the applicable models at a uniform 5 seconds likewise leaves the calibration picture intact [2].
Classifier-free guidance as a dial
Because Cosmos3-Super was the one model run locally rather than through a hosted API, the authors could sweep its guidance scale; the other five expose no adjustable guidance parameter through their public endpoints [2]. On the dice scene, with a shorter 61-frame horizon used throughout the sweep, the trade-off is direct:
| Guidance scale | Scorability | Total variation | p-value |
|---|---|---|---|
| 1.0 | 0.66 | 0.17 | 0.699 |
| 1.5 | 0.78 | 0.17 | 0.394 |
| 3.0 | 0.97 | 0.34 | 0.003 |
| 4.5 | 1.00 | 0.26 | 0.042 |
| 6.0 (default) | 0.97 | 0.31 | 0.004 |
| 7.5 | 1.00 | 0.45 | below 0.001 |
| 9.0 | 1.00 | 0.51 | below 0.001 |
Turning guidance down to 1.0 or 1.5 produces a distribution statistically indistinguishable from uniform, but scorability falls to 0.66 and 0.78 because the model stops following the prompt well enough to yield readable outcomes. At guidance of 3.0 and above scorability saturates while the distribution sharpens onto faces 1 and 5, with face 6 essentially absent [2]. A separate ablation asked each model to roll a named face and measured compliance against the 16.7 percent chance baseline: SeeDance-2.0 complied 63.3 percent of the time, WAN-2.7 46.4 percent, Veo 3.1 43.3 percent, Cosmos3-Super 33.3 percent, Runway Gen-4.5 28.6 percent, and HappyHorse exactly at chance [2].
The models evaluated
All six systems named in the benchmark are real, publicly released models rather than internal or hypothetical baselines:
| Model | Evidence |
|---|---|
| WAN-2.7 | Alibaba's Wan 2.7 image-to-video model, served on Replicate as wan-video/wan-2.7-i2v [7] |
| SeeDance-2.0 | ByteDance Seed's Seedance 2.0, released in China in early February 2026 and documented in arXiv:2604.14148 [9] |
| HappyHorse-1.0 | Alibaba's Happy Horse 1.0 text-to-video and image-to-video model, served on Replicate as alibaba/happyhorse-1.0 [6] |
| Veo 3.1 | Google DeepMind's Veo 3.1 [10] |
| Runway Gen-4.5 | Announced by Runway on December 1, 2025 [8] |
| Cosmos3-Super | A tier of NVIDIA's Cosmos 3 family (arXiv:2606.02800), published on Hugging Face as nvidia/Cosmos3-Super [11][12] |
Odyssey's own world models, including Odyssey-3 and Starchild-1, are not part of the evaluated set [2]. That is worth stating plainly in both directions: CaliBench is authored by a company that builds and sells competing world models, so its leaderboard is not a neutral third-party ranking, but the published results also contain no self-favourable entry, and the committed outcome data lets anyone recompute the tables.
Relationship to other video benchmarks
CaliBench is positioned against two existing evaluation families. Distributional metrics such as Frechet Video Distance, introduced by Unterthiner and colleagues in 2018, compare whole datasets inside a learned perceptual feature space [4]. The paper's objection is that such distances are uninterpretable in physical terms and aggregate over a whole dataset, so they cannot isolate the uncertainty of one specific phenomenon. CaliBench instead scores in a discrete, physically meaningful space (a bin index, a die face, a suit, a colour) where distance from a known reference can be computed directly [2].
The second family is per-sample evaluation. VBench scores video generations along more than sixteen dimensions and includes a diversity measure that detects near-identical outputs, but the paper argues sample diversity is not the same property as calibration, since a diverse model can still be badly mis-distributed [2][5]. Benchmarks such as PhysicsIQ, which compares a generated continuation against a single ground-truth trajectory, and PhyGenBench, which uses a vision-language judge for physical common sense, both operate at the level of the individual video [2]. Odyssey's suggested usage is additive: report mnTV alongside FVD and VBench scores rather than instead of them [1].
Release and availability
The protocol was released on GitHub at odysseyml/CaliBench on the day of the announcement [1][3]. The repository contains the nine scene definitions (conditioning frame stem plus generation prompt) in scenes.py, the reference distributions and per-scene outcome counts in constants.py, the conditioning frames themselves, the vision-language extraction prompts, the analysis scripts that produce the paper's tables and figures, and the mnTV implementation [3].
Notably, the extracted outcomes for every scene-model cell are committed to the repository as JSON, so the paper's numbers can be recomputed with no API calls at all; only the generation and extraction stages need paid Replicate and Gemini keys [3]. The repository also ships the statistical machinery separately: a bootstrap confidence interval script for mnTV, a Monte Carlo chi-squared script, a power analysis at N = 32, and a sensitivity check that brackets the effect of imputing the unscorable generations [3]. GitHub reported no license file for the repository as of mid-September 2026, so reuse terms are unstated.
The paper's arXiv listing records acceptance at Transactions on Machine Learning Research and presentation at the WOOP workshop at ECCV 2026 [2]. Odyssey lists CaliBench on its research page alongside Odyssey-3, Agora-1 and PROWL-1 [13].
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10Odyssey, "Introducing CaliBench: Are World Models Physically Calibrated?", August 10, 2026. odyssey.systems/introducing-calibench
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33 ^34 ^35J. Sadeghi, J. Seidenschwarz, J. Allardice, S. Srinivasan, B. Graham, J. Hawke, "CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?", arXiv:2608.16829. arxiv.org/...2608.16829
- ^1 ^2 ^3 ^4 ^5 ^6odysseyml/CaliBench, GitHub repository. github.com/...CaliBench
- ^1 ^2T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, S. Gelly, "Towards Accurate Generative Models of Video: A New Metric and Challenges", arXiv:1812.01717. arxiv.org/...1812.01717
- ^Z. Huang et al., "VBench: Comprehensive Benchmark Suite for Video Generative Models", arXiv:2311.17982. arxiv.org/...2311.17982
- ^Replicate, "alibaba/happyhorse-1.0". replicate.com/...happyhorse-1.0
- ^Replicate, "wan-video/wan-2.7-i2v". replicate.com/...wan-2.7-i2v
- ^Runway, "Runway Gen-4.5: State-of-the-Art AI Video Generation", December 1, 2025. runwayml.com/...introducing-runway-gen-4.5
- ^ByteDance Seed, "Seedance 2.0: Advancing Video Generation for World Complexity", arXiv:2604.14148. arxiv.org/...2604.14148
- ^Google DeepMind, "Veo 3.1". deepmind.google/...veo
- ^N. Agarwal et al., "Cosmos 3: Omnimodal World Models for Physical AI", arXiv:2606.02800. arxiv.org/...2606.02800
- ^Hugging Face, "nvidia/Cosmos3-Super". huggingface.co/...Cosmos3-Super
- ^Odyssey, "Research". odyssey.systems/research
- ^Google, "Gemini 3.1 Pro". blog.google/...gemini-3-1-pro
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 3,088 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Adversarial verification 2026-09-16 (cluster V3): every figure checked against arXiv 2608.16829 v2 and the CaliBench repo; 1 minor defect fixed
Cite this page: AI Wiki. "CaliBench." aiwiki.ai, updated 16 Sept 2026, fact-checked 16 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/calibench