Physics-IQ Verified
Physics-IQ Verified is an audited and re-scored version of Physics-IQ, the Google DeepMind benchmark that tests whether video generation models can predict how real physical experiments unfold. It was introduced in the paper "Physics-IQ Verified" (arXiv:2606.18943), submitted on 17 June 2026 by Tim Rädsch, Yuki M. Asano, Hilde Kuehne, Stefan Bauer, Priyank Jaini, Robert Geirhos and Carsten T. Lüth.[1] The audit rewrote unclear text prompts, removed recording artifacts from the ground-truth activation masks, and replaced the original dataset-level score with a per-sample score that weights every sample and metric equally. According to the authors, the changes refine 57.6% of all samples and improve over 34.8% of prompts, and in a six-model study they moved the ranking enough to give a Kendall's τ of 0.46 between the old and new orderings.[1] A public leaderboard for the benchmark is run by Anates Labs, a Munich research lab with which both of the paper's joint leads are affiliated, and a combined ranking table is also kept in the DeepMind GitHub repository, which now calls Physics-IQ Verified the "recommended benchmark variant".[3][5] As of 30 September 2026 the top two image-to-video entries were fine-tuned NVIDIA Cosmos 3 models from the Physis-Lang project, at 48.2% and 43.3%.[3]
Overview
| Item | Detail |
|---|---|
| Type | Benchmark for physical understanding in video generation models (audited version of Physics-IQ)[1] |
| Paper | "Physics-IQ Verified", arXiv:2606.18943 (cs.CV), v1 submitted 17 June 2026[1] |
| Authors | Tim Rädsch, Yuki M. Asano, Hilde Kuehne, Stefan Bauer, Priyank Jaini, Robert Geirhos, Carsten T. Lüth (Rädsch and Lüth are joint leads)[2] |
| Base benchmark | Physics-IQ (Motamed et al., WACV 2026): 66 experiments, 396 real videos[2][5] |
| Evaluation set | 198 take-1 videos (66 scenarios × 3 camera angles); 198 second takes used to measure physical variation[2] |
| Metrics | Spatial IoU, Spatiotemporal IoU, Weighted Spatial IoU, MSE, combined into a per-sample Physics-IQ Verified score[2] |
| Tracks | Image-to-video (I2V) and video-to-video / multiframe (V2V)[3][5] |
| Code | github.com/google-deepmind/physics-IQ-benchmark (verified workflow is the default)[5] |
| Data | Hugging Face dataset Anates-Labs-Research/Physics-IQ-Verified, CC BY 4.0, gated with automatic approval[6] |
| Leaderboard | physics-iq-verified.anates.ai (Anates Labs) and the GitHub README[3][5] |
Background: the original Physics-IQ
Physics-IQ was introduced by Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini and Robert Geirhos in "Do generative video models understand physical principles?" (arXiv:2501.09038, first submitted 14 January 2025, published at WACV 2026).[8][5] It covers fluid dynamics, optics, solid mechanics, magnetism and thermodynamics, and its authors reported that for the models they tested (Sora, Runway, Pika, Lumiere, Stable Video Diffusion and VideoPoet) physical understanding was "severely limited, and unrelated to visual realism".[8]
The benchmark contains 66 experiments, each filmed from three angles and performed twice, for 66 × 3 × 2 = 396 videos recorded at 3840 × 2160 and 30 frames per second.[2][5] Each 8-second clip is split into a 3-second conditioning segment and a 5-second continuation. For the first 198 videos a "switch frame" marks the 3-second point where generation should start; the second 198 videos are repeat takes, used to measure how much two real runs of the same experiment differ.[2] A text-to-video model receives only a prompt, an image-to-video model gets the switch frame plus the prompt, and a video-to-video model gets the conditioning clip plus the prompt.[2]
Scoring compares the generated continuation with the real one using four metrics. Three are intersection-over-union (IoU) scores computed on binary activation maps derived from frame-to-frame differences, and one is pixel-level mean squared error:[2]
| Metric | Question it answers (paper's wording) |
|---|---|
| Spatial IoU | Where does action happen? |
| Spatiotemporal IoU | Where and when does action happen? |
| Weighted spatial IoU | Where and how much does action happen? |
| Mean squared error (MSE) | How does action happen? |
In the original formulation, each metric is averaged over the dataset and divided by the "physical variation" (the same metric computed between the two real takes). The IoU terms are then averaged and the MSE term is subtracted.[2] The Verified paper notes that Physics-IQ has been widely reported by later model papers, including MAGI-1, Sora 2 and physics-alignment work, and was the target of an ICCV 2025 Physics-IQ Challenge. The authors argue that this adoption makes measurement errors in the benchmark more consequential, because they can "propagate into model-development decisions".[2]
The audit: what was wrong
The paper identifies three sources of measurement error: unclear prompts, spurious activations ("artifacts") in the ground truth, and the way scores are aggregated.[2]
Unclear prompts
The authors define a good prompt as one that works like an exam question. A person given the prompt and the start frame should be able to predict the outcome with high confidence, but the prompt must not give the answer away.[2] They sorted flawed prompts into four categories, in decreasing order of severity:[2]
- Factually incorrect: the prompt does not match what happens in the video.
- Temporally imprecise: the prompt does not separate actions that already happened before the switch frame from those the model should generate.
- Omitted key information: the prompt lacks information needed to model the effect, such as an object's material.
- Vague language: the prompt describes the action too loosely to constrain generation.
The first two make a correct generation impossible in principle; the last two leave physical degrees of freedom open and increase output variance.[2] The Anates Labs dataset-fixes page gives worked examples. In scenario 0179 the original prompt describes "A blue and yellow tennis ball" when the balls in the recording are grey and brown, and the page says models responded by spawning extra blue and yellow balls. In scenario 0170 the prompt describes a basket being lowered over a ball when the event to be predicted is the basket lifting. In scenario 0125 the prompt says "A yellow mug" without saying it is ceramic, so it is unclear whether a mug that fails to break reflects poor physics or missing information.[4]
Artifacts in the ground truth
Because the IoU metrics compare activation maps from the real and generated videos, anything that moves in the real recording counts as part of the target. The paper defines an artifact as "a metric activation caused by a visual event that is not part of the physical effect under observation" and splits artifacts into two types:[2]
- Deterministic artifacts come from events that the prompt or setup specifies, such as a rotating platform or stick. They are predictable, but their activation reflects the apparatus rather than the physical effect under study.
- Non-deterministic artifacts arise by chance during recording, for example a grabber tool moving after it releases an object. No model or human can anticipate them, so the paper calls them "more damaging from a measurement perspective".
Aggregation
The original score is defined only over the whole dataset. Because the physical-variation denominator is averaged across all samples, the paper argues that experiments with low natural variation are down-weighted (they can never reach a score of 1) while those with high variation are up-weighted (their individual scores can exceed 1). A dataset-level score also hides which samples a model fails on.[2] An appendix adds that the unclipped IoU subscores are unbounded, so one exceptional subscore can dominate the composite, which the authors say contradicts the original authors' intent that "no metric should be assessed in isolation".[2]
How common the problems were
Across the 198 evaluation videos, the paper counts 69 videos with unclear prompts and 59 with artifacts, with 20 in both groups.[2] Its summary figure and the dataset-fixes page break these down as follows (the categories can overlap):
| Issue | Paper (share of 198 videos)[2] | Dataset-fixes page (cases)[4] |
|---|---|---|
| Unclear prompts (any) | 34.8% | n/a |
| Factually incorrect | 4.5% | 12 |
| Temporally imprecise | 4.5% | 12 |
| Omitted key information | 27.3% | 54 |
| Vague language | 4.5% | 9 |
| Artifacts (any) | 29.8% | n/a |
| Deterministic artifacts | 6.1% | 12 |
| Non-deterministic artifacts | 23.7% | 47 |
| Videos modified overall | 57.6% | n/a |
The two sources agree on omitted information, vague language and both artifact types, but the dataset-fixes page lists 12 cases each for factually incorrect and temporally imprecise prompts, while the paper's figure gives 4.5% (about nine videos) for each.[2][4] At the frame level, the paper reports that 66.2% of frames contain at least one activation, and that 27.1% of those active frames were modified by artifact removal.[2]
The three fixes
1. Corrected and templated prompts
The authors made "minimally invasive corrections" to the unclear prompts, then split every prompt into six fields so that model-specific "templaters" could assemble the prompt each model's provider recommends.[2]
| Field | Type | Content |
|---|---|---|
| SETUP | Variable | Pre-action scene: objects, arrangement, initial conditions |
| SCENE | Variable | Temporally constant scene information |
| ACTION | Variable | Subject-action description |
| CAM | Fixed | "Static locked-off single-shot with fixed frame throughout, filmed at constant framerate in real-time." |
| STYLE (new) | Fixed | Constrains output to "a realistic scientific demonstration" |
| SCOPE (new) | Fixed | States that the scene "only contains the described setup and actions" |
The original prompts ended with "Static shot with no camera movement". The rewrite expresses every instruction in positive terms, citing research showing that language and vision models handle negation poorly and provider guidance that discourages it.[2] The paper's appendix shows the effect: with the original prompt, Wan 2.2 generated a hand that interacted with a stationary rubber duck, and P-Video zoomed the camera in on a rotating teapot, while the templated prompts suppressed both behaviours.[2] Prompts built this way are called best-practice prompts (bpp); the originals are labelled op. The repository ships a base template plus model-specific templates for P-Video and Sora 2, and lets submitters register their own templater.[5]
2. Artifact cleaning
Artifacts were removed from the ground truth by manual annotation and "frame freezing", which holds pixel values constant in an affected region from a given time onward. The authors chose freezing over masking or inpainting because it adds no new visual information and creates no artificial boundaries that could themselves trigger activations.[2] Two annotation types drive it:[2]
end_effect_framesmarks when the physical effect ends; everything after it is frozen.freeze_areasmarks regions and start times for artifacts that occur during the effect but away from it, such as a moving grabber or a still-spinning rotator.
For some deterministic-artifact scenarios the improved prompt was also changed so that the rotating base stops once the effect has been set in motion. The fixes page shows this for the domino scenarios 0057 and 0051, whose verified prompts end "Then the rotation stops."[2][4]
3. Sample-level scoring
Physics-IQ Verified computes a score for each sample. Each IoU metric is divided by that sample's own physical variation and clipped to [0, 1]. MSE is turned into a "higher is better" term by dividing the physical-variation MSE by the generated MSE and clipping. The four terms are averaged with equal weight, and the benchmark score is the mean over all samples.[2] The authors say the main benefit is traceability: a low score can be traced back to specific samples and failure modes.[2]
Ranking-change study
The paper evaluated six image-to-video models: three open-source (Wan 2.2, HunyuanVideo 1.5, Cosmos3-Nano) and three closed-source (Sora 2, P-Video from Pruna AI, Grok Imagine Video). Each model generated four full sets of 198 videos with both prompt sets, and results were compared in a factorial design: 2 prompt sets × 2 ground truths × 2 scores.[2] The paper lists the generation settings and May 2026 per-video prices it used:[2]
| Model | Size | FPS | Resolution | Seed control | Price per video |
|---|---|---|---|---|---|
| Grok Imagine Video | n.d. | 24 | 1280×720 | No | $0.352 |
| HunyuanV-1.5 | 8.3B | 24 | 848×480 | Yes | $0.400 |
| P-Video | n.d. | 24 | 1280×704 | Yes | $0.100 |
| Sora 2 | n.d. | 30 | 1280×720 | No | $0.800 |
| Wan 2.2 | 14B | 16 | 1280×720 | Yes | $0.110 |
| Cosmos3-Nano | 16B | 24 | 1280×720 | Yes | $0.333 |
Comparing the fully original evaluation (original prompts, original ground truth, original score) with the fully verified one (best-practice prompts, cleaned ground truth, verified score) gave these results:[2]
| Model | Original Physics-IQ | Original rank | Physics-IQ Verified | Verified rank |
|---|---|---|---|---|
| Wan 2.2 | 35.4 ± 1.2 | 1 | 32.2 ± 0.6 | 3 |
| Grok Imagine Video | 32.9 ± 0.4 | 2 | 34.8 ± 0.6 | 1 |
| HunyuanV-1.5 | 29.7 ± 1.0 | 3 | 33.4 ± 0.8 | 2 |
| P-Video | 22.5 ± 2.0 | 4 | 25.3 ± 1.8 | 6 |
| Cosmos3-Nano | 21.7 ± 1.9 | 5 | 29.1 ± 2.4 | 4 |
| Sora 2 | 12.7 ± 0.8 | 6 | 26.5 ± 0.8 | 5 |
Scores are means ± standard deviation over four runs. Wan 2.2 was the only model whose score fell.[2] The rank correlations between the two orderings were Spearman's ρ = 0.65 and Kendall's τ = 0.46, which the authors describe as "moderate but meaningful" changes. A bootstrap analysis over 500 resampled video sets gave mean cross-evaluation correlations of ρ̄ = 0.697 and τ̄ = 0.513, while correlations within each evaluation exceeded 0.9, and the 95% confidence intervals did not overlap.[2]
The paper then isolated each change:[2]
- Prompts. Best-practice prompts gave significantly better subscores across all primary metrics (Wilcoxon signed-rank, all p < 0.05; Cohen's d ≥ 0.55). Sora 2 benefited most, largely because the new prompts cut unwanted camera motion; its original score rose from 12.7 to 25.3. Wan 2.2 was the only model that got worse with the new prompts.
- Artifact removal. Cleaning the ground truth significantly reduced all IoU-based scores and the original composite for every model (p ≪ 10⁻⁵, Cohen's d ≤ -1), with the largest drop for Wan 2.2. The authors read this as a sign that some of Wan 2.2's lead came from confounding effects. For the spatiotemporal metric, physical variation rose by about 17% under the verified protocol, so the paper warns that scores should not be compared across protocols without adjusting for that difference.
- Scoring. The sample-level score raised every model's number but did not change the ranking; bootstrap correlations between the two scoring schemes were close to 1.
Two further findings bear on how the leaderboard is read. First, the authors found that Sora 2 performed "notably worse in April 2026 than in October 2025": a single October 2025 run with original prompts scored 40.6 on the verified ground truth, against 15.7 for the April 2026 runs.[2] Second, they passed prompts to Cosmos3-Nano without any language-model rewriting, noting that the official Cosmos 3 leaderboard score used upsampled prompts, and reported that upsampling best-practice prompts with Opus 4.8 added about one point in separate experiments.[2]
Benchmark release and workflow
The verified benchmark lives in the existing DeepMind repository rather than a separate project. The README describes Physics-IQ Verified as the "recommended benchmark variant", and verified evaluation is the default behaviour of physiq/run_physics_iq.py; the original benchmark is evaluated by adding --original_physics_iq.[5] The verified videos, switch frames and masks are downloaded from the Hugging Face dataset Anates-Labs-Research/Physics-IQ-Verified, which is licensed CC BY 4.0 and gated with automatic approval.[5][6] The repository states that its software is licensed under Apache 2.0 and other materials under CC BY 4.0, and that it "is not an official Google product". Its submission section also carries a disclaimer that Physics-IQ Verified "is an independent third-party benchmark that is not endorsed or verified by Google DeepMind".[5]
The workflow has submitters choose or write a prompt templater, generate videos (from switch frames for I2V, or from 3-second conditioning clips for V2V), trim every output to exactly 5 seconds, and run the evaluator, which reports both the original and the verified score for each run folder.[5] The leaderboard rules in the README are:[5]
- One run is enough to be listed, but four runs reporting mean and standard deviation are recommended, and four-run standard deviations are required to claim state of the art on either track.
- Entries must link a technical report, reproducible repository or other documentation of the method.
- Because the original and verified benchmarks share scenarios, signal from one must not be used to improve scores on the other.
Unless an entry is marked op, all leaderboard scores use best-practice prompts.[5] New entries arrive as pull requests to the DeepMind repository. The ones merged in September 2026 were merged by Robert Geirhos.[10][11][12][13]
The leaderboard
Anates Labs hosts the leaderboard at physics-iq-verified.anates.ai. It shows a "Verified Ranking" for either the I2V or V2V track, lets viewers filter by company, open-source or proprietary availability, and an "LLM Usage" flag, and can rank models or companies.[3] A "Net Improvement" view shows each score relative to the average of the entries in the selected track.[3] A metric breakdown lists the leaders for each submetric (Spatial, Spatiotemporal, Weighted Spatial and MSE), and a cost frontier plots score against generation cost in three views: native cost, FPS-normalized cost, and cost normalized for both FPS and resolution.[3] The site's footnote reads: "Price via leading API providers or estimated via GPU market rate, May 2026. Generation cost is normalized to 24 FPS and 1280-wide output. Separate LLM prompt overhead is added after video normalization where used."[3] A companion page, "Dataset-fix visual overview", shows original and verified prompts, switch frames and accumulated activation heatmaps for the fixes.[4]
I2V ranking as of 30 September 2026
The table reproduces the site's I2V Verified Ranking as fetched on 30 September 2026. Company, availability, LLM flag and date are shown as the site lists them. Normalized cost is the site's FPS-and-resolution-normalized figure, where one is given.[3]
| Rank | Model | Company (as listed) | Availability | Physics-IQ Verified | LLM | Date | Normalized cost |
|---|---|---|---|---|---|---|---|
| 1 | Physis-Lang (Cosmos3 Super) | NVIDIA | Open source | 48.2% ± 1.4 | Yes | 2026-09-25 | n/a |
| 2 | Physis-Lang (Cosmos3 Nano) | NVIDIA | Open source | 43.3% ± 1.5 | Yes | 2026-09-25 | n/a |
| 3 | Cosmos3 Super | NVIDIA | Open source | 42.7% ± 0.8 | Yes | 2026-09-04 | $0.823 |
| 4 | Seedance 2.5 | Seedance | Proprietary | 42.4% ± 0.8 | Yes | 2026-09-28 | $2.838 |
| 5 | MiniMax H3 | MiniMax | Open source | 39.8% ± 0.3 | Yes | 2026-08-24 | $0.438 |
| 6 | Cosmos3 Nano | NVIDIA | Open source | 37.3% ± 0.9 | Yes | 2026-09-04 | $0.434 |
| 7 | MiniMax H3 Max | fal.ai | Proprietary | 36.2% ± 0.7 | Yes | 2026-08-27 | $0.482 |
| 8 | Gemini Omni 1.1 Flash | Proprietary | 35.3% ± 0.4 | Yes | 2026-09-28 | $0.609 | |
| 9 | Grok Imagine Video | xAI | Proprietary | 34.8% ± 0.6 | Yes | 2026-06-17 | $0.352 |
| 10 | Magi-1 24B + GeoPhys (BoN) (op) | Sand AI | Open source | 33.7% ± 1.4 | Yes | 2026-06-19 | n/a |
| 11 | Hunyuan Video 1.5 | Tencent | Open source | 33.4% ± 0.8 | No | 2026-06-17 | $0.604 |
| 12 | Gemini Omni Flash | Proprietary | 33.4% ± 0.2 | Yes | 2026-09-28 | $0.609 | |
| 13 | Cosmos3 Edge | NVIDIA | Open source | 32.7% ± 1.1 | Yes | 2026-08-31 | n/a |
| 14 | Wan 2.2 14B | Alibaba | Open source | 32.2% ± 0.6 | No | 2026-06-17 | $0.165 |
| 15 | Veo 3.1 Lite | Proprietary | 31.8% ± 0.3 | Yes | 2026-09-28 | $0.300 | |
| 16 | CogVideoX-5B | Z.ai | Open source | 31.8% ± 1.5 | No | 2026-08-18 | $1.256 |
| 17 | Kandinsky-WM 1.0 | Kandinsky Lab | Open source | 30.8% ± 0.9 | Yes | 2026-08-07 | $0.222 |
| 18 | Veo 3.1 Fast | Proprietary | 30.0% ± 0.5 | Yes | 2026-09-28 | $0.600 | |
| 19 | Wan 2.2 5B | Alibaba | Open source | 27.7% ± 0.9 | No | 2026-08-18 | $0.090 |
| 20 | Sora 2 | OpenAI | Proprietary | 26.5% ± 0.8 | Yes | 2026-06-17 | $0.640 |
| 21 | P-Video | Pruna AI | Proprietary | 25.3% ± 1.8 | No | 2026-06-17 | $0.100 |
The five entries dated 2026-06-17 (Grok Imagine Video, Hunyuan Video 1.5, Wan 2.2 14B, Sora 2 and P-Video) carry the same scores the paper reports for its verified evaluation.[2][3] Cosmos3-Nano does not: the paper's run without prompt rewriting scored 29.1, while the leaderboard entry uses LLM-upsampled prompts (see below).[2][3]
For the submetrics, the site listed Physis-Lang (Cosmos3 Super) first in all four as of 30 September 2026: Spatial 59.9 ± 1.4 (ahead of MiniMax H3 at 58.9), Spatiotemporal 41.6 ± 3.1 (ahead of CogVideoX-5B at 35.5), Weighted Spatial 48.5 ± 1.1 and MSE 43.0 ± 1.3.[3] On cost, Seedance 2.5 had the highest normalized cost of the priced entries ($2.838), and Wan 2.2 5B the lowest ($0.090).[3]
The dates in the GitHub README do not always match the site. The README lists both Physis-Lang entries as added on 2026-09-28, the Seedance, Gemini and Veo entries on 2026-09-27, and MiniMax H3 on 2026-09-04.[5] The README table also includes a Magi-1 24B (op) I2V entry at 30.2 that the site's I2V ranking does not show.[3][5]
V2V ranking as of 30 September 2026
The multiframe (V2V) track had four entries on 30 September 2026:[3][5]
| Rank | Model | Company | Physics-IQ Verified | Date (site) |
|---|---|---|---|---|
| 1 | Magi-1 24B + GeoPhys (BoN) (op) | Sand AI | 58.2% ± 1.8 | 2026-06-19 |
| 2 | Cosmos3 Super V2V | NVIDIA | 50.8% ± 2.2 | 2026-09-18 |
| 3 | Magi-1 24B (op) | Sand AI | 48.4% ± 1.1 | 2026-06-19 |
| 4 | Cosmos3 Nano V2V | NVIDIA | 43.0% ± 2.0 | 2026-09-18 |
The NVIDIA V2V results were submitted in pull request 82, opened on 7 September 2026 and merged on 22 September.[12] "BoN" marks a best-of-N sampling protocol, and "op" marks original prompts.[3][5]
Notable leaderboard events
Cosmos 3 score revision (August-September 2026)
On 31 August 2026 a member of NVIDIA's Cosmos team opened GitHub issue 75, reporting that the leaderboard's figures for Cosmos3-Nano (30.3 ± 0.6) and Cosmos3-Super-Image2Video (39.5 ± 0.8) were lower than the team's own reproductions.[9] Using best-practice prompts upsampled with Opus 4.8 (following Cosmos documentation) and four seeds, the Cosmos team member measured 42.68 ± 0.76 for Super and 37.25 ± 0.89 for Nano. They also measured 34.39 ± 1.20 for Nano with raw best-practice prompts and no upsampling, which still beat the published figure.[9] A second user reported better-than-published results for Nano without upsampling.[9] The maintainers updated the entries to 42.7 ± 0.8 and 37.3 ± 0.9 in pull request 81, merged on 4 September 2026.[10] With that revision, Cosmos3 Super held the top I2V score of 42.7 until the Physis-Lang entries were merged on 28 September.[5][10][11]
Late-September additions
Pull request 85, opened on 27 September and merged on 28 September 2026, added Seedance 2.5 (42.43 ± 0.76), Gemini Omni 1.1 Flash (35.34 ± 0.44), Gemini Omni Flash Preview (33.36 ± 0.19), Veo 3.1 Lite (31.83 ± 0.32) and Veo 3.1 Fast (29.96 ± 0.53). The pull request describes them as "vanilla runs" using unmodified base best-practice prompts, four runs each, at native 24 FPS with 120 evaluated frames.[13]
Pull request 83, opened on 25 September and merged on 28 September 2026, added Physis-Lang (Cosmos3 Nano) at 43.29 ± 1.52 and Physis-Lang (Cosmos3 Super) at 48.23 ± 1.43.[11] The submitters stated that these "are fine-tuned Cosmos-based models, not the unmodified Cosmos baselines", that the results "are self-reported and submitted for verification", and that the runs used the team's own prompts ("base captions plus physical reasoning") with the Cosmos default negative prompt, not the official best-practice prompts. They reported four seeds per model, 198 videos per run, 1280×720 output at 24 FPS, and no best-of-N selection.[11] The Physis-Lang paper, from authors at NVIDIA, MIT and the University of Oxford, gives different Physics-IQ Verified figures from its own evaluation: Cosmos3-Nano rising from 40.23 to 43.41 and Cosmos3-Super from 45.92 to 50.00 after Physis-Lang fine-tuning.[14]
Limitations
The paper's own caveats include the fixed set of 198 evaluation videos, the reliance on manual artifact annotations, the evaluation of only six image-to-video models, and the fact that reference-based scoring can still penalise continuations that are physically plausible but differ from the recorded take.[2] Its related-work section adds that Physics-IQ-style mask metrics do not test whether quantities such as energy or momentum are conserved.[2]
Leaderboard entries are submitted by model developers or third parties, and the rules allow a single run.[5] Several entries use language-model prompt rewriting or best-of-N sampling, which the site flags with its LLM column and entry labels, so rows are not all produced under the same protocol.[3][5] The Cosmos 3 revision showed that inference settings and prompt upsampling can move a model's score by several points.[9][10] The paper's Sora 2 comparison also shows that a hosted model's score can change over time without a change in its name.[2]
Anates Labs
Anates Labs describes itself on its website as working on "Visual intelligence from first principles", and lists Physics-IQ Verified and "Visual intelligence" as its research areas.[7] The site's legal notice names the operator as Anates Space Labs GmbH in Munich, represented by Dr. Tim Raedsch.[7] In the paper, Tim Rädsch (Anates Labs and the Technical University of Munich) and Carsten T. Lüth (Anates Labs) are listed as joint leads. The other authors are affiliated with the University of Technology Nuremberg (Asano), the Tübingen AI Center at the University of Tübingen (Kuehne), the Technical University of Munich and Helmholtz AI in Munich (Bauer), and Google DeepMind (Jaini and Geirhos). Jaini and Geirhos, both co-authors of the original Physics-IQ paper, "contributed in an advisory capacity". The acknowledgments thank Pruna AI for model credits.[2][8]
See also
References
- ^1 ^2 ^3 ^4Rädsch, T.; Asano, Y. M.; Kuehne, H.; Bauer, S.; Jaini, P.; Geirhos, R.; Lüth, C. T. "Physics-IQ Verified". arXiv:2606.18943, submitted 17 June 2026. arxiv.org/...2606.18943
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33 ^34 ^35 ^36 ^37 ^38 ^39 ^40 ^41 ^42 ^43Rädsch, T. et al. "Physics-IQ Verified" (full text, v1 PDF, including appendices). arXiv, 2026. arxiv.org/...2606.18943
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17Anates Labs. "Physics-IQ Verified" leaderboard (I2V and V2V Verified Ranking, metric breakdown, cost frontier). Retrieved 30 September 2026. physics-iq-verified.anates.ai
- ^1 ^2 ^3 ^4 ^5Anates Labs. "Dataset-fix visual overview", Physics-IQ Verified. Retrieved 30 September 2026. physics-iq-verified.anates.ai/dataset-fixes
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21Google DeepMind. "google-deepmind/physics-IQ-benchmark" (README: "Physics-IQ and Physics-IQ Verified: Benchmarking physical understanding in generative video models"). GitHub. Retrieved 30 September 2026. github.com/...physics-IQ-benchmark
- ^1 ^2Anates Labs Research. "Anates-Labs-Research/Physics-IQ-Verified" dataset. Hugging Face. Retrieved 30 September 2026. huggingface.co/...Physics-IQ-Verified
- ^1 ^2Anates Labs. "Anates Labs: Visual intelligence from first principles" (including legal notice). Retrieved 30 September 2026. anates.ai
- ^1 ^2 ^3Motamed, S.; Culp, L.; Swersky, K.; Jaini, P.; Geirhos, R. "Do generative video models understand physical principles?". arXiv:2501.09038, 2025. arxiv.org/...2501.09038
- ^1 ^2 ^3 ^4"Cosmos3 Nano & Super-I2V IQ-Verified Score Reproduction Discrepancies". Issue 75, google-deepmind/physics-IQ-benchmark, GitHub, opened 31 August 2026. github.com/...75
- ^1 ^2 ^3 ^4"Update Cosmos3 Physics-IQ Verified scores". Pull request 81, google-deepmind/physics-IQ-benchmark, GitHub, merged 4 September 2026. github.com/...81
- ^1 ^2 ^3 ^4"Add Physis-Lang (Cosmos3 Nano) and Physis-Lang (Cosmos3 Super) to Physics-IQ Verified". Pull request 83, google-deepmind/physics-IQ-benchmark, GitHub, merged 28 September 2026. github.com/...83
- ^1 ^2"Cosmos3 V2V (multiframe) results for Physics-IQ Verified Leaderboard". Pull request 82, google-deepmind/physics-IQ-benchmark, GitHub, merged 22 September 2026. github.com/...82
- ^1 ^2"Add Seedance, Gemini Omni, and Veo 3.1 verified leaderboard entries". Pull request 85, google-deepmind/physics-IQ-benchmark, GitHub, merged 28 September 2026. github.com/...85
- ^Lu, L.; Ma, X.; He, W.; Zhan, G. et al. "Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model". Paper PDF, Physis-Intelligence/Physis-Lang, GitHub, 2026. github.com/...Physis-Lang_arxiv.pdf
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,746 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent verification (xg15 V6, 30 Sep 2026): leaderboard snapshot rows, paper tables, audit figures, GitHub PRs and HF dataset checked; 1 minor defect fixed in v2
Cite this page: AI Wiki. "Physics-IQ Verified." aiwiki.ai, updated 30 Sept 2026, fact-checked 30 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/physics_iq_verified