# Physics-IQ

> Source: https://aiwiki.ai/wiki/physics_iq
> Updated: 2026-09-30
> Fact-checked: 2026-09-30
> Categories: AI Benchmarks, Google DeepMind, Video Generation, World Models
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Physics-IQ." aiwiki.ai, 30 Sept 2026. https://aiwiki.ai/wiki/physics_iq
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

**Physics-IQ** is a benchmark that tests whether generative video models understand physical principles, by asking them to continue real videos of physical experiments and comparing the result with what actually happened. It was introduced by Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini and Robert Geirhos in the paper "Do generative video models understand physical principles?" (arXiv:2501.09038), first posted on 14 January 2025 and published at the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026.[1][4] Motamed is listed with INSAIT, Sofia University (work done while at [Google DeepMind](https://aiwiki.ai/wiki/google_deepmind)); the other four authors are listed with Google DeepMind, and the code and data are released from the google-deepmind GitHub organization.[2][6] The dataset holds 396 real videos of 66 experiments covering solid mechanics, fluid dynamics, optics, thermodynamics and magnetism.[2] Its headline finding was that every model tested ([Sora](https://aiwiki.ai/wiki/sora), [Runway Gen 3](https://aiwiki.ai/wiki/runway_gen_3), [Pika](https://aiwiki.ai/wiki/pika) 1.0, [Lumiere](https://aiwiki.ai/wiki/lumiere), [Stable Video Diffusion](https://aiwiki.ai/wiki/stable_video_diffusion) and VideoPoet) scored far below the level set by natural variation between two real recordings, with the best at 29.5 out of 100, and that how realistic a model's videos looked was not significantly correlated with how well it followed physics. In the authors' words, "visual realism does not imply physical understanding".[1][2]

Physics-IQ became a common reporting target for video and [world model](https://aiwiki.ai/wiki/world_model) developers, including Sand.ai's MAGI-1, Meta researchers' WMReward and [NVIDIA Cosmos 3](https://aiwiki.ai/wiki/nvidia_cosmos_3), and it was run as a guest challenge at the Perception Test Workshop at [ICCV](https://aiwiki.ai/wiki/iccv) 2025.[9][10][12][13] In June 2026 an audit by a group that included two of the original authors produced [Physics-IQ Verified](https://aiwiki.ai/wiki/physics_iq_verified), which the GitHub repository now calls the "recommended benchmark variant".[6][14]

## Overview

| Item | Detail |
|---|---|
| Type | Benchmark for physical understanding in generative video models[1] |
| Paper | "Do generative video models understand physical principles?", arXiv:2501.09038 (v1 14 January 2025, v3 27 February 2025)[1] |
| Venue | WACV 2026, pages 948-958[4] |
| Authors | Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, Robert Geirhos (Jaini and Geirhos joint last authors)[2] |
| Affiliations (paper) | INSAIT, Sofia University (Motamed, work done while at Google DeepMind); Google DeepMind (others)[2] |
| Data | 396 videos: 66 scenarios × 3 camera angles × 2 takes, 8 seconds each, 3840 × 2160 at 30 fps[2] |
| Categories | Solid mechanics, fluid dynamics, optics, thermodynamics, magnetism[2] |
| Tracks | Image-to-video (i2v, one "switch frame") and multiframe or video-to-video (v2v, up to 3 seconds of conditioning video)[2][6] |
| Metrics | Spatial IoU, Spatiotemporal IoU, Weighted spatial IoU, MSE, combined into the Physics-IQ score[2] |
| Score scale | Normalized so that two real takes of the same experiment score 100[2] |
| Best score in the paper | 29.5 (VideoPoet, multiframe)[2] |
| Code and data | github.com/google-deepmind/physics-IQ-benchmark; software under Apache 2.0, other materials under CC BY 4.0[6] |
| Successor | [Physics-IQ Verified](https://aiwiki.ai/wiki/physics_iq_verified) (arXiv:2606.18943, June 2026)[14] |

## Background

The paper frames the benchmark around a debate about [AI video generation](https://aiwiki.ai/wiki/ai_video_generation): whether models trained to predict how videos continue learn "world models" that discover the laws of physics, or whether they are "sophisticated pixel predictors" that reach visual realism without understanding.[2] The authors set out both sides. Proponents argue that next-frame prediction forces a model to learn trajectories, gravity and fluid behavior, by analogy with next-token prediction in language models and with predictive processing in the brain. Skeptics argue that watching video is passive, so a model cannot observe the effects of its own interventions and must separate correlation from causation without them; on this view, realistic output can come from reproducing common patterns in training data.[2]

The authors also argued that most earlier physical-reasoning benchmarks were a poor fit for this question. Physion, Physion++, CRAFT, IntPhys, CoPhy, CLEVRER, ESPRIT and PhyWorld use synthetic scenes, which introduces a real-versus-synthetic distribution shift for models trained on natural video, while VideoPhy and [PhyGenBench](https://aiwiki.ai/wiki/phygenbench) assess physical commonsense from text prompts rather than from real footage.[2] Physics-IQ instead uses real recordings and asks models to do what they are trained for: predict how a video continues.[2]

## Dataset

Each of the 66 scenarios targets a specific physical principle. The paper lists collisions, object continuity, occlusion, object permanence, fluid dynamics, chain reactions, trajectories under forces such as gravity, material properties and reactions, and lights, shadows, reflections and magnetism.[2] Examples in the paper include a chain of dominoes that falls either normally or with a rubber duck placed in the middle of the chain, and pillows onto which either a kettlebell or a piece of paper is dropped.[2]

Every scenario was filmed from three fixed perspectives (left, center and right) with Sony Alpha a6400 cameras and 16-50 mm lenses, at 3840 × 2160 and 30 frames per second, with no camera motion. Each scenario was shot twice under identical conditions. The authors call the difference between the two takes "physical variance", which captures chaotic motion, small changes in friction and variation in how forces were applied.[2] That gives 66 × 3 × 2 = 396 videos, each 8 seconds long.[2] The repository distributes the videos at 30, 24, 16 and 8 fps; it generates other frame rates from the 30 fps files when needed.[6]

## Evaluation protocol

Each 8-second video is split into a 3-second conditioning clip and a 5-second test clip that serves as ground truth.[2] Models get one of two kinds of conditioning:

- **Image-to-video (i2v).** The model receives the last frame of the conditioning clip, called the switch frame. The authors chose each switch frame by hand so that it shows enough of the setup while leaving the outcome to be predicted; in the domino scenario, for example, it is the moment the first domino has tipped but has not yet touched the second.[2]
- **Multiframe (video-to-video, v2v).** Models that accept several frames receive as much of the 3-second conditioning clip as they can take.[2]

Both kinds of model can also receive a human-written description of the conditioning part of the scene, written so that it does not give away what happens next. Stable Video Diffusion was the only model in the original study that does not accept text.[2] The model then generates a 5-second continuation. For each model, the authors resampled the Physics-IQ videos to the model's preferred frame rate (8 to 30 fps across the models tested) and resolution.[2] The current repository requires generated videos to be trimmed to exactly 5 seconds before scoring.[6] Generations are made for the 198 take-1 videos (66 scenarios × 3 views); take 1 is the reference continuation, and take 2 is compared with take 1 to estimate physical variation.[6][15]

## Metrics and the Physics-IQ score

The authors argued that standard video-quality metrics (PSNR, SSIM, FVD and LPIPS) measure appearance and feature statistics, not whether motion and interactions are physically correct.[2] Physics-IQ therefore uses four metrics. Because the camera never moves, a threshold on pixel changes between frames (with background subtraction and morphological cleaning) produces a binary "motion mask video" marking where movement occurs in each frame.[2]

| Metric | Question (paper's heading) | How it is computed |
|---|---|---|
| Spatial IoU | Where does action happen? | Collapse the motion mask over time with a max operation into a binary map of everywhere motion occurred, then take the [intersection over union](https://aiwiki.ai/wiki/iou) with the real video's map.[2] |
| Spatiotemporal IoU | Where and when does action happen? | Compare the two motion mask videos frame by frame and average over time, so the timing of motion also counts.[2] |
| Weighted spatial IoU | Where and how much action happens? | Collapse the motion mask by averaging per-frame action rather than taking the max, then sum the pixel-wise minimum of the two maps and divide by the sum of the pixel-wise maximum. This separates a pendulum swinging repeatedly through an area from a ball passing once.[2] |
| [Mean squared error](https://aiwiki.ai/wiki/mean_squared_error_mse) | How does action happen? | Average squared pixel difference between real and generated frames; lower is better. It penalizes changes in appearance such as an object changing color.[2] |

The four metrics are combined into the Physics-IQ score by summing them, with MSE entering with a negative sign, and the score is normalized so that physical variance, which the authors call "the upper limit of what we can reasonably expect a model to capture", equals 100%.[2] The Physics-IQ Verified paper writes out the original aggregation: each metric is averaged over the whole dataset and divided by the same metric computed between the two real takes (for MSE, the real-versus-real value is subtracted instead); the three IoU ratios are averaged, the MSE term is subtracted, and the result is clipped to the range 0 to 1.[15] The paper reports a Spearman correlation of -0.92 (p < .005) between this aggregate score and each model's mean rank across the four metrics, and argues that collapsing to one number largely preserves the ranking.[2] The authors state that none of the metrics measures a physical quantity directly; they are proxies, and "no metric should be assessed in isolation".[2]

## Visual realism test

To measure how realistic the generated videos look, separate from whether they follow physics, the authors used a two-alternative forced-choice test with a multimodal language model, [Gemini 1.5 Pro](https://aiwiki.ai/wiki/gemini_1_5_pro). For each scenario the model saw a real and a generated video in random order and was asked which one was generated. Accuracy near the 50% chance level means the generated videos were hard to tell apart from real ones.[2] The prompt, given in the paper's appendix, framed the task as helping to sort a mix of real camera footage and computer-generated videos, and included an offer to "tip you $100" for good work, before asking the model to explain its reasoning and finish with a fixed answer statement.[2]

Gemini identified the generated video up to 86.9% of the time (for Lumiere multiframe). Sora was hardest to detect, with an MLLM score of 55.6%, followed by Runway Gen 3 at 74.8% and VideoPoet (multiframe) at 77.3%.[2] The authors noted that Gemini's stated reasons for its choices were often unrelated to the video content, which they compared with post-hoc rationalization in human experiments.[2] Plotting the MLLM score against the Physics-IQ score gave a Pearson correlation of r = -0.46 with p = .249, which the paper reports as not significant.[2]

## Results of the original study

The published version evaluates eight model configurations: VideoPoet and Lumiere in both i2v and multiframe modes, and Runway Gen 3, Pika 1.0, Stable Video Diffusion and Sora in i2v mode only.[2] Luma was left out because, according to the paper, its usage policy prohibited benchmarking, and [Veo 2](https://aiwiki.ai/wiki/veo_2) was not generally available when the paper was written.[2]

| Model (mode) | Spatial IoU ↑ | Spatiotemporal IoU ↑ | Weighted spatial IoU ↑ | MSE ↓ | Physics-IQ score ↑ |
|---|---|---|---|---|---|
| Physical variance (two real takes) | 0.678 | 0.535 | 0.577 | 0.002 | 100.0 |
| VideoPoet (multiframe) | 0.204 | 0.164 | 0.137 | 0.010 | 29.5 |
| Lumiere (multiframe) | 0.170 | 0.155 | 0.093 | 0.013 | 23.0 |
| Runway Gen 3 (i2v) | 0.201 | 0.115 | 0.116 | 0.015 | 22.8 |
| VideoPoet (i2v) | 0.141 | 0.126 | 0.087 | 0.012 | 20.3 |
| Lumiere (i2v) | 0.113 | 0.173 | 0.061 | 0.016 | 19.0 |
| Stable Video Diffusion (i2v) | 0.132 | 0.076 | 0.073 | 0.021 | 14.8 |
| Pika 1.0 (i2v) | 0.140 | 0.041 | 0.078 | 0.014 | 13.0 |
| Sora (i2v) | 0.138 | 0.047 | 0.063 | 0.030 | 10.0 |

*Table 1 of arXiv:2501.09038v3.[2]*

The main findings reported by the authors:

- **A large gap to real videos.** The best configuration, VideoPoet (multiframe), reached 29.5 against the physical-variance ceiling of 100. The authors noted that VideoPoet is a causal model.[2]
- **More context helps.** For the two models run in both modes, the multiframe version beat the i2v version, as expected when temporal information is available.[2]
- **Location is easier than timing and amount.** All models did much better on Spatial IoU than on the stricter metrics. No physics category was "solved", though performance varied by category.[2] The project page adds that fluid dynamics tended to work better than solid mechanics.[5]
- **Some scenarios were already solved.** VideoPoet (multiframe) plausibly continued paint being smeared on glass, and Runway Gen 3 did so for red liquid poured over a rubber duck, while both failed on a ball falling into a crate and on a tangerine being cut with a knife.[2]
- **Sora's low score partly reflects shot changes.** The authors described Sora's videos as often "visually and artistically superior" but noted that they frequently contained transition cuts despite instructions not to change the camera view, which several metrics penalize. They expected that a version following the static-camera prompt more closely would score "substantially" higher.[2]
- **Hallucination and dataset bias.** In one scenario a burning match is lowered into water; Runway Gen 3 instead generated a candle that appeared and was lit by the match, with each frame realistic but the sequence physically impossible. In prototyping, Lumiere turned a red pool table green as soon as it began generating, which the authors read as a bias toward the more common green tables.[2]

The authors did not conclude that video models cannot learn physics. They wrote that it remains open whether scaling next-frame prediction will close the gap or whether more interactive training will be needed, that they were "optimistic" about future-frame prediction, and that inference-time scaling, "such as sampling more", might also improve results.[2]

## Versions and publication

| Date | Version | Notes |
|---|---|---|
| 14 January 2025 | arXiv v1 | Titled "Do generative video models learn physical principles from watching videos?"; best score 24.1 (VideoPoet multiframe), Sora 8.7[1][3] |
| 10 February 2025 | arXiv v2 | Same title and headline figures as v1[1][19] |
| 17 February 2025 | Code fix and leaderboard update | A commit "fixed bug discovered by running take-2 videos", and a same-day commit "update readme with correct leaderboard results" replaced the README scores (VideoPoet multiframe 24.1 to 29.5, Sora 8.7 to 10.0)[7][8] |
| 27 February 2025 | arXiv v3 | Retitled "Do generative video models understand physical principles?", with the revised scores[1][2] |
| March 2026 | WACV 2026 | Published in the conference proceedings, pages 948-958[4] |

The v1 and v3 tables differ for every model, and the order below first place also changed: in v1, Runway Gen 3 (18.4) ranked second and Lumiere multiframe (18.2) third, while in v3 Lumiere multiframe (23.0) is second and Runway Gen 3 (22.8) third.[3][2][8] Figures quoted from the first version, such as a best score of 24.1%, are therefore out of date.

## Code, data and license

The repository google-deepmind/physics-IQ-benchmark contains the download scripts, the evaluation code (`physiq/run_physics_iq.py`), the text descriptions and the leaderboards.[6] The original dataset is hosted in a Google Cloud Storage bucket; the evaluator is run with the `--original_physics_iq` flag for the original benchmark, and without it for Physics-IQ Verified, which is now the default.[6] According to the README, all software is licensed under Apache 2.0 and all other materials under Creative Commons Attribution 4.0 (CC BY); the README also states that it "is not an official Google product".[6] The README describes the ceiling this way: the best possible score is 100.0%, "achieved by physically realistic videos that differ only in physical randomness but adhere to all tested principles of physics".[6]

The leaderboard is open to outside submissions: developers who test a model can open a pull request adding a row.[6] Leaderboard rows are therefore self-reported by the submitting team, and entries differ in prompting and inference setup.

## Adoption

### Model reports

- **MAGI-1 (Sand.ai, May 2025).** The MAGI-1 technical report called Physics-IQ "the most comprehensive and state-of-the-art benchmark" in its area and reported 56.02 in the video-to-video setting (using the full 24 fps, 96-frame conditioning clip) and 30.23 in image-to-video, which the report said was about 27 points above VideoPoet in v2v. The report attributed the result to MAGI-1's autoregressive design.[10] See [Sand.ai](https://aiwiki.ai/wiki/sand_ai).
- **WMReward (Meta and collaborators, 2025-2026).** Researchers from FAIR at Meta, the University of Oxford, Mila and other institutions used [V-JEPA 2](https://aiwiki.ai/wiki/v_jepa_2) as a reward signal to steer and rerank MAGI-1 generations. A short technical report describes this as the winning entry of the Physics-IQ Challenge at ICCV 2025.[11] The full paper (arXiv:2601.10553) applies the method, named WMReward, with [best-of-N sampling](https://aiwiki.ai/wiki/best_of_n_sampling) over 16 candidates, and gives the official challenge-platform results as 37.39 for MAGI-1 i2v and 62.64 for v2v.[12] It also applies the reward to Wan 2.2 and, through its API, to [Sora 2](https://aiwiki.ai/wiki/sora_2) (best-of-N only, since guidance is unavailable for an API model), raising Sora 2 from 42.30 to 46.4 in i2v. In the paper's i2v comparison, best-of-N selection judged by the vision-language model Qwen2.5-VL lowered the score of all four base models it was tried on, while best-of-N with the V-JEPA 2 reward raised them by about 4 to 7 points.[12]
- **VLM-guided self-refinement (University of Chinese Academy of Sciences, November 2025).** A training-free method in which a vision-language model (Gemini 2.5 Pro) and a language model (GPT-4o) iteratively rewrite MAGI-1's prompt based on detected physics errors reported a Physics-IQ score of 62.38, against a 56.31 baseline. That figure comes from an ensemble that combines the best outputs of six refinement runs; each single refinement loop scored between 48.31 and 52.92, below the baseline.[16]
- **NVIDIA Cosmos 3 (2026).** The Cosmos 3 technical report uses Physics-IQ as its dedicated test of physics adherence, because other benchmarks treat physical plausibility as only one scoring component. It reports Cosmos3-Super at 43.8 (i2v) and 59.7 (v2v), rising to 48.9 and 63.4 with WMReward best-of-N reranking, and describes both as state of the art. For these runs NVIDIA generated text prompts with its own prompt upsampler rather than using the benchmark's descriptions.[13]

### Original leaderboard as of 30 September 2026

The repository's "Physics-IQ Original Leaderboard" mixes plain model results with results that use inference-time search. Selected entries, with their leaderboard ranks:[6]

| Rank | Entry | Input | Score | Date added |
|---|---|---|---|---|
| 1 | Magi-1 + GeoPhys (BoN) | v2v | 64.5% | 2026-06-17 |
| 2 | Cosmos3-Super + WMReward (BoN) | v2v | 63.4% | 2026-05-26 |
| 3 | Magi-1 + WMReward (BoN) | v2v | 62.6% | 2025-10-28 |
| 4 | Cosmos3-Super | v2v | 59.7% | 2026-05-26 |
| 5 | Cosmos3-Nano + WMReward (BoN) | v2v | 57.7% | 2026-05-26 |
| 6 | Magi-1 | v2v | 56.0% | 2025-04-21 |
| 8 | Cosmos3-Super + WMReward (BoN) | i2v | 48.9% | 2026-05-26 |
| 9 | Sora2 + WMReward (BoN) | i2v | 46.4% | 2026-04-01 |
| 11 | Cosmos3-Super | i2v | 43.8% | 2026-05-26 |
| 13 | Sora2 | i2v | 42.3% | 2026-04-01 |
| 25 | VideoPoet (from the original paper) | v2v | 29.5% | 2025-02-19 |
| 32 | Sora (from the original paper) | i2v | 10.0% | 2025-02-19 |

"BoN" marks best-of-N selection among several generated candidates using a learned reward, so those entries spend extra inference compute per test video. The Sora 2 entries cite the WMReward paper as their source.[6][12]

### ICCV 2025 Physics-IQ Challenge

Physics-IQ was a guest challenge at the Perception Test Workshop at ICCV 2025 in Honolulu, organized by Robert Geirhos and Priyank Jaini (Google DeepMind) with Luc Van Gool and Saman Motamed (INSAIT, Sofia University).[9] The Eval.ai submission server opened on 1 August 2025, submissions were due on 2 October, and prizes were €1,000, €750 and €500 for the top three teams, which had to submit short technical reports. The workshop's Physics-IQ track included a keynote by Phillip Isola.[9] The winning entry was the V-JEPA 2 reward method described above.[11][12] INSAIT said in October 2025 that the benchmark had "attracted wide attention across the global AI research community" following its presentation at ICCV 2025.[17]

## Limitations and criticism

The original authors described several limits of their own design. The metrics are proxies rather than direct physical measurements, and they may be "on the conservative side" because they strongly penalize hallucinated objects, camera movement and shot changes; the authors called this "not ideal" but argued that scientific benchmarks "should err on the side of caution".[2] The MLLM realism test is limited by the capability of the model used as judge.[2] They also observed that Stable Video Diffusion produced many hallucinations and implausible motions yet scored in the same Spatial IoU range as Lumiere, Sora, Pika and VideoPoet (i2v).[2]

The Physics-IQ Verified paper added further criticisms:[15]

- **Aggregation.** Because the original score is defined only over the whole dataset, it cannot show which samples a model fails on. Averaging physical variation across the dataset gives low-variation experiments less weight and lets high-variation experiments exceed a ratio of 1. The unclipped IoU ratios also let one exceptional sub-score dominate the composite.
- **Ground-truth artifacts.** Movement unrelated to the physical event, such as the grabber tools used to hold and release objects or rotating platforms, can trigger motion-mask activations in the reference videos.
- **Prompts.** Some original text descriptions were factually wrong, badly timed relative to the switch frame, or missing key information.
- **Scope.** Overlap metrics can confuse spatial proximity with physical correctness and assume that the recorded trajectory is the only correct outcome, so a physically plausible alternative can be scored as a failure. The benchmark does not check whether quantities such as energy or momentum are conserved, and the models in the original study have since been superseded.

The same paper argued that Physics-IQ's adoption as "a standard reporting protocol and an optimization target" makes these measurement errors more consequential, because they can "propagate into model-development decisions".[15]

## Physics-IQ Verified

[Physics-IQ Verified](https://aiwiki.ai/wiki/physics_iq_verified) (arXiv:2606.18943, submitted 17 June 2026) is an audit and revision of the benchmark by Tim Rädsch, Yuki M. Asano, Hilde Kuehne, Stefan Bauer, Priyank Jaini, Robert Geirhos and Carsten T. Lüth. It keeps the real-world continuation setup but corrects prompts, removes artifact-driven activations from the ground truth, and replaces the dataset-level score with a sample-level score that weights each sample and metric equally. The authors report that it refines 57.6% of samples and improves over 34.8% of prompts, and that re-scoring six image-to-video models changed their ranking moderately (Kendall's τ = 0.46).[14] The Verified workflow is now the default in the google-deepmind repository.[6] The README also carries a disclaimer that Physics-IQ Verified "is an independent third-party benchmark that is not endorsed or verified by Google DeepMind".[6] A public leaderboard is hosted by Anates Labs; as of 30 September 2026 its top image-to-video entry was [Physis-Lang](https://aiwiki.ai/wiki/physis_lang) (Cosmos3 Super) at 48.2% ± 1.4.[18] The Verified paper cautions that scores should not be compared across the original and Verified protocols without accounting for the difference in measured physical variance.[15]

## See also

- [Physics-IQ Verified](https://aiwiki.ai/wiki/physics_iq_verified)
- [PhyGenBench](https://aiwiki.ai/wiki/phygenbench)
- [Physis-Lang](https://aiwiki.ai/wiki/physis_lang)
- [World model](https://aiwiki.ai/wiki/world_model)
- [AI video generation](https://aiwiki.ai/wiki/ai_video_generation)
- [NVIDIA Cosmos 3](https://aiwiki.ai/wiki/nvidia_cosmos_3)
- [V-JEPA 2](https://aiwiki.ai/wiki/v_jepa_2)
- [Inference-time scaling](https://aiwiki.ai/wiki/inference_time_scaling)

## References

1. Motamed, S.; Culp, L.; Swersky, K.; Jaini, P.; Geirhos, R. "Do generative video models understand physical principles?" arXiv:2501.09038 (abstract page and submission history: v1 14 Jan 2025, v2 10 Feb 2025, v3 27 Feb 2025). https://arxiv.org/abs/2501.09038
2. Motamed, S. et al. "Do generative video models understand physical principles?" arXiv:2501.09038v3, full text (HTML), 27 February 2025. https://arxiv.org/html/2501.09038v3
3. Motamed, S. et al. "Do generative video models learn physical principles from watching videos?" arXiv:2501.09038v1, full text (HTML), 14 January 2025. https://arxiv.org/html/2501.09038v1
4. Motamed, S.; Culp, L.; Swersky, K.; Jaini, P.; Geirhos, R. "Do Generative Video Models Understand Physical Principles?" Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026, pp. 948-958. CVF Open Access. https://openaccess.thecvf.com/content/WACV2026/html/Motamed_Do_Generative_Video_Models_Understand_Physical_Principles_WACV_2026_paper.html
5. "Physics IQ Benchmark: Do generative video models understand physical principles?" Project page. Retrieved 30 September 2026. https://physics-iq.github.io/
6. Google DeepMind. "google-deepmind/physics-IQ-benchmark" (README: "Physics-IQ and Physics-IQ Verified: Benchmarking physical understanding in generative video models"; leaderboards, workflows, license). GitHub. Retrieved 30 September 2026. https://github.com/google-deepmind/physics-IQ-benchmark
7. google-deepmind/physics-IQ-benchmark. Commit 8241d41, "fixed bug discovered by running take-2 videos + Robert's added video lebgth checks and ID-only names for generated videos", 17 February 2025. https://github.com/google-deepmind/physics-IQ-benchmark/commit/8241d417e5b739e7d16d7edd468d26a9e054837d
8. google-deepmind/physics-IQ-benchmark. Commit 0b041b8, "update readme with correct leaderboard results", 17 February 2025. https://github.com/google-deepmind/physics-IQ-benchmark/commit/0b041b84d1e7fb9d9752776175045ce49d276d8f
9. "Physics-IQ Challenge @ ICCV 2025: Physics-IQ × Perception Test Challenge". Retrieved 30 September 2026. https://physics-iq.github.io/workshop/physics-iq-challenge.html
10. Sand.ai; Teng, H. et al. "MAGI-1: Autoregressive Video Generation at Scale". arXiv:2505.13211, 19 May 2025. https://arxiv.org/abs/2505.13211
11. Yuan, J.; Zhang, X.; Friedrich, F. et al. "Improving the Physics of Video Generation with VJEPA-2 Reward Signal". arXiv:2510.21840, 22 October 2025. https://arxiv.org/abs/2510.21840
12. Yuan, J.; Zhang, X.; Friedrich, F. et al. "Inference-time Physics Alignment of Video Generative Models with Latent World Models". arXiv:2601.10553 (v1 15 Jan 2026, v2 27 Feb 2026). https://arxiv.org/abs/2601.10553
13. NVIDIA. "Cosmos 3: Omnimodal World Models for Physical AI" (technical report, Table 13 and Physics-IQ section). Retrieved 30 September 2026. https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdf
14. Rädsch, T.; Asano, Y. M.; Kuehne, H.; Bauer, S.; Jaini, P.; Geirhos, R.; Lüth, C. T. "Physics-IQ Verified". arXiv:2606.18943, submitted 17 June 2026. https://arxiv.org/abs/2606.18943
15. Rädsch, T. et al. "Physics-IQ Verified", full text (HTML v1, including appendix C on the original score aggregation). arXiv, 2026. https://arxiv.org/html/2606.18943v1
16. Liu, Y.; Zhao, X.; Wen, P.; Dai, S.; Huang, Q. "Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement". arXiv:2511.20280, 25 November 2025. https://arxiv.org/abs/2511.20280
17. INSAIT. "New Benchmark 'Physics-IQ' Challenges AI Video Models' Understanding of the Physical World". 21 October 2025. https://insait.ai/new-benchmark-physics-iq-challenges-ai-video-models-understanding-of-the-physical-world/
18. Anates Labs. "Physics-IQ Verified" leaderboard. Retrieved 30 September 2026. https://physics-iq-verified.anates.ai/
19. Motamed, S. et al. "Do generative video models learn physical principles from watching videos?" arXiv:2501.09038v2, full text (HTML), 10 February 2025. https://arxiv.org/html/2501.09038v2

