PhyGenBench
PhyGenBench (Physics Generation Benchmark) is a benchmark for testing whether text-to-video models produce videos that follow basic physical commonsense. It was introduced in the paper "Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation", first posted to arXiv on 7 October 2024 by Fanqing Meng, Jiaqi Liao and colleagues from OpenGVLab at the Shanghai AI Laboratory, Shanghai Jiao Tong University, the University of Hong Kong and the Chinese University of Hong Kong, and later published at ICML 2025.[1][2] The benchmark has 160 text prompts covering 27 physical laws in four domains: mechanics, optics, thermal phenomena and material properties.[1] It ships with an automatic evaluation framework, PhyGenEval, which uses GPT-4o to turn each prompt into physics-specific questions and then scores the generated video in three stages with vision-language models.[1][3] In the original evaluation, no model scored above 0.51 on a 0-1 scale in the October 2024 arXiv version, or above 0.55 in the ICML version, which the authors read as evidence that current video generators are far from being "world simulators".[1][2] By 2026 PhyGenBench had become one of the standard physics benchmarks reported in video generation papers, including NVIDIA's Physis-Lang, which re-scored it with GPT-5.5 as the judge.[6]
| Field | Value |
|---|---|
| Full name | PhyGenBench (Physics Generation Benchmark), with the PhyGenEval evaluation framework |
| Paper | "Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation"[1] |
| First release | arXiv:2410.05363, submitted 7 October 2024 (labelled "Technical report")[1] |
| Venue | Proceedings of the 42nd International Conference on Machine Learning (ICML 2025), PMLR 267:43781-43806[2] |
| Developers | OpenGVLab, Shanghai AI Laboratory; Shanghai Jiao Tong University; University of Hong Kong; Chinese University of Hong Kong[1] |
| Corresponding authors | Wenqi Shao and Ping Luo[1][2] |
| Task | Text-to-video generation |
| Size | 160 prompts, 27 physical laws, 4 domains[1] |
| Scoring | Physical commonsense alignment (PCA) on a 0-3 scale per video, reported normalized to 0-1; separate semantic alignment (SA) score[1][2] |
| Original judges | GPT-4o, VQAScore, CLIP retrieval, LLaVA-Interleave, InternVideo2[1][3] |
| Code and data | github.com/OpenGVLab/PhyGenBench[3] |
| Project page | phygenbench123.github.io[4] |
Background
The paper starts from the argument, often made after OpenAI's Sora, that large text-to-video models could become general simulators of the physical world. The authors cite cognitive psychology research on "intuitive physics", the physical expectations that even human infants show, and argue that a world model should first reproduce simple, everyday physical phenomena correctly.[1] One of the paper's examples is a generated video in which a stone placed on water fails to sink.[1]
Existing video benchmarks, the authors wrote, mostly measured visual quality (motion smoothness, background consistency) or spatial relationships, and reference-based metrics such as FVD need a ground-truth video that does not exist for new scenes. VideoPhy (Bansal et al., 2024) did target physical commonsense, but the PhyGenBench authors argued its prompts were not tied to explicit physical laws.[1] VideoPhy has 688 prompts built around solid-solid, solid-fluid and fluid-fluid interactions; PhyGenBench's appendix argues that these prompts are often too short or describe processes that cannot be seen (such as a watch spring winding inside a case), so generators fail on semantics before physics can be judged. In a 64-prompt comparison, human-rated semantic alignment averaged 0.63 on VideoPhy prompts and 0.80 on PhyGenBench prompts across three models.[1]
Benchmark design
The authors define physical commonsense as a basic intuitive understanding of how objects and actions behave in everyday life, and divide it into four areas following Halliday, Resnick and Walker's textbook Fundamentals of Physics. Each prompt is meant to show one clear physical phenomenon tied to one law, so that annotators and automatic judges are not confused by several laws acting at once.[1]
| Domain | Laws or phenomena covered | Prompts | Example prompt (from the paper) |
|---|---|---|---|
| Mechanics | 7: gravity, buoyancy, solid pressure, atmospheric pressure, elasticity, friction, surface tension | 40 | "A piece of iron is gently placed on the surface of the water in a tank filled with water" (the iron should sink) |
| Optics | 6: reflection, refraction, scattering, dispersion, interference and diffraction, straight-line propagation | 50 | "a kite soaring above a smooth and tranquil pond" (the kite should be reflected) |
| Thermal | 6 phase transitions: solidification, melting, liquefaction, boiling, deposition, sublimation | 30 | a timelapse of water as the temperature rises above 100 degrees Celsius |
| Material properties | 5 physical properties (color, hardness, solubility, combustibility, flame reaction) and 3 chemical properties (acidity, redox potential, dehydrating properties) | 40 | an egg hurled with force at a rock (the egg should break and the rock stay intact) |
The prompts contain 165 unique objects and 42 unique actions, with an average length of 18.75 words.[1] Other examples shown in the paper include juice poured out on a space station (liquid should float in globules), a copper flame (should be green), sulfuric acid poured on bread (the bread should char through dehydration) and ice cream heated above 100 degrees Celsius.[1]
Construction
Prompts were built in five steps. The authors (1) picked laws from textbooks that are widely recognized and easy to observe, (2) hand-wrote an initial prompt for each law, (3) expanded prompts with extra detail about objects and actions, because modern generators are trained on long, descriptive captions, taking care not to state the expected outcome, (4) used GPT-4o to swap objects (for example egg to vase, glass bottle or cup; rock to wall or metal) for diversity, following T2V-CompBench, and (5) reviewed the prompts and the associated laws manually, and spot-checked that current models could generate the scenes semantically.[1] The GitHub release contains the prompts (prompts.json), an "explicit captions" file used for the prompt-rewriting experiment, and example questions for each PhyGenEval stage.[3]
PhyGenEval
PhyGenEval splits evaluation into two scores. Semantic alignment (SA) asks whether the video shows the objects and action in the prompt. Physical commonsense alignment (PCA) asks whether the expected physical outcome actually happens. Both are graded on a four-point scale (0-3), matching the human rating scale.[1]
For SA, GPT-4o first extracts the objects and action from the prompt, then checks the video for the objects (0-2) and for the action (0 or 1).[1]
For PCA, the authors argued that asking a video-understanding model directly whether a clip is physically correct does not work, because such models understand physics poorly. PhyGenEval instead first uses GPT-4o, working only on text, to analyse the prompt and its physical law and write stage-specific questions, then evaluates the video in three increasingly broad stages.[1]
| Stage | What it checks | How | Judge models |
|---|---|---|---|
| 1. Key physical phenomena detection | Whether the key outcome (for example "the egg breaks") appears at the right moment | A retrieval prompt locates the key frame (the paper follows the CLIP-based CLIPScore method); questions are scored on that frame and its four neighbours; affirmative and negative statements are compared | VQAScore (single image)[1][3] |
| 2. Physics order verification | Whether events happen in the causal order (the egg touches the rock, then breaks) | Three multi-image questions: first frame to key frame, key frame to last frame, and the whole sequence | GPT-4o or LLaVA-Interleave-DPO-7B[1][3] |
| 3. Overall naturalness evaluation | Whether the whole clip looks physically realistic | GPT-4o writes a prompt-specific grading rubric from a general rubric adapted from DEVIL; the video model chooses one of four levels: Completely Fantastical, Clearly Unrealistic, Slightly Unrealistic, Almost Realistic | GPT-4o or InternVideo2[1][3] |
Each stage is discretized to 0-3; the three stage scores are averaged and rounded down. The authors recommend an ensemble that runs stage 2 with both GPT-4o and LLaVA-Interleave and stage 3 with both GPT-4o and InternVideo2.[1] The README also offers a closed-source-only setting in which only VQAScore needs to be installed locally, and says all evaluations ran on a single A100-80G GPU.[3] The README gives the GPT-4o snapshot as "gpt4o-0513", while a footnote in the ICML paper gives "gpt4o-0806".[2][3]
Agreement with human ratings
For human evaluation, three annotators rated 64 randomly chosen prompts for each of eight models (512 videos) on the same 0-3 SA and PCA scales.[1] Correlation between PhyGenEval's PCA and these human scores was far higher than for earlier automatic metrics:[1][2]
| Metric | Overall Kendall's tau | Overall Spearman's rho |
|---|---|---|
| DEVIL (Gemini 1.5 Pro judge) | 0.17 | 0.18 |
| VideoPhy (VideoCon-Physics) | 0.03 | 0.04 |
| VideoScore | 0.17 | 0.19 |
| PhyGenEval | 0.78 | 0.81 |
Ablations in the arXiv version reported that each of the three stages on its own reached a Spearman correlation between 0.46 and 0.61, and that only the combination reached 0.81. Using only open-source judges gave 0.66, and using only GPT-4o gave 0.71.[1] Case studies showed the older metrics rating an egg that bounced like rubber, or a stone floating on water, as physically correct.[1]
Original results
The arXiv technical report evaluated eight models: five open-source (CogVideoX 2B and 5B, Open-Sora 1.2, LaVie, Vchitect 2.0) and three commercial (Kling, Pika and Runway Gen-3), generating 1,280 videos.[1] The ICML camera-ready expanded this to 14 models and 2,240 videos, adding HunyuanVideo, two Pyramid Flow variants, Sora, Vidu and Hailuo.[2]
PCA scores from the ICML version (normalized to 0-1, higher is better):[2]
| Model | Size | Mechanics | Optics | Thermal | Material | Average |
|---|---|---|---|---|---|---|
| CogVideoX | 2B | 0.38 | 0.43 | 0.34 | 0.39 | 0.39 |
| CogVideoX | 5B | 0.39 | 0.55 | 0.40 | 0.42 | 0.45 |
| Open-Sora 1.2 | 1.1B | 0.43 | 0.50 | 0.44 | 0.37 | 0.44 |
| LaVie | 860M | 0.30 | 0.44 | 0.38 | 0.32 | 0.36 |
| Vchitect 2.0 | 2B | 0.41 | 0.56 | 0.44 | 0.37 | 0.45 |
| HunyuanVideo | 13B | 0.44 | 0.53 | 0.38 | 0.39 | 0.47 |
| Pyramid Flow (flux) | 2B | 0.37 | 0.50 | 0.47 | 0.37 | 0.43 |
| Pyramid Flow (sd3) | 2B | 0.42 | 0.49 | 0.36 | 0.45 | 0.41 |
| Pika 1.0 | n/a | 0.35 | 0.56 | 0.43 | 0.39 | 0.44 |
| Runway Gen-3 | n/a | 0.45 | 0.57 | 0.49 | 0.51 | 0.51 |
| Kling | n/a | 0.45 | 0.58 | 0.50 | 0.40 | 0.49 |
| Sora | n/a | 0.50 | 0.66 | 0.56 | 0.46 | 0.55 |
| Vidu | n/a | 0.48 | 0.63 | 0.52 | 0.50 | 0.54 |
| Hailuo | n/a | 0.49 | 0.67 | 0.50 | 0.50 | 0.55 |
The authors' main observations were that every model scored best on optics, which they attributed to optical phenomena being common and explicit in pretraining data; that the best commercial models led clearly; and that LaVie showed consistently low physical correctness, with the lowest average (0.36), although CogVideoX 2B scored below it in optics and thermal.[1][2] Qualitatively, all models failed to show a glass ball sinking in a bathtub, most models got melting wrong (one CogVideoX video showed ice cream growing), Kling showed the egg bouncing off the rock like a rubber ball, and no model got the sulfuric-acid-on-bread reaction right.[1] Semantic alignment was much higher than physical alignment: machine SA averages ran from 0.58 (LaVie) to 0.85 (Gen-3 and Kling), which the authors took as evidence that the scenes themselves were easy to render.[1]
Some text in the ICML version was carried over from the arXiv report without being updated. Its introduction and results section still say "even the best-performing model, Gen-3, scores only 0.51", although its own Table 2 lists Sora and Hailuo at 0.55 and Vidu at 0.54.[2] The leaderboard in the GitHub README, which predates the ICML version, also differs from the paper's table in a few cells: it lists CogVideoX 2B's average as 0.37 (paper: 0.39), CogVideoX 5B's mechanics score as 0.43 (paper: 0.39) and LaVie's mechanics score as 0.40 (paper: 0.30).[1][3]
Experiments on possible fixes
The paper tested three common ways to improve generation:[1][2]
| Approach | Setup | Result |
|---|---|---|
| Scaling | CogVideoX 2B versus 5B | Average PCA rose from 0.39 to 0.45, with better static scenes (soap-bubble interference, rust on iron) but little gain in mechanics; neither model rendered a bouncing ball correctly |
| Prompt rewriting | GPT rewrote prompts to state the expected outcome (for example, "The liquid forms floating globules") | CogVideoX 5B rose from 0.45 to 0.52 and Kling from 0.49 to 0.56; fixed simple cases such as flame color but not the egg breaking or the stone sinking |
| Video enhancement | Vchitect 2.0 videos passed through VEnhancer | Average unchanged at 0.45, although VEnhancer raised Vchitect's VBench score above Kling's; model scores before and after had a Spearman correlation of 0.86 |
The authors concluded that scaling and prompt engineering do not solve dynamic physical scenarios, suggested training on curated synthetic data, and argued that the VEnhancer result shows PhyGenEval measures physical correctness separately from visual quality.[1]
Versions and release
The arXiv version (v1, 7 October 2024, the only version as of September 2026) lists ten authors: Fanqing Meng, Jiaqi Liao, Xinyu Tan, Wenqi Shao, Quanfeng Lu, Kaipeng Zhang, Yu Cheng, Dianqi Li, Yu Qiao and Ping Luo, with Meng and Liao as equal first authors.[1] The ICML 2025 proceedings version lists nine authors in a different order (Meng, Liao, Tan, Lu, Shao, Zhang, Cheng, Li, Luo), without Yu Qiao.[2][5] Code and data were released on GitHub on 7 October 2024, along with a project page of example videos.[3][4] As of September 2026 the repository describes itself as "[ICML2025] The code and data" of the paper and does not declare a software license.[3]
Adoption
Research papers
PhyGenBench became a common evaluation for methods that try to make video generators more physically plausible. Several of these papers reuse the baseline numbers from the original table and add their own model; PhysRAG and VideoCoCo state that they follow the official protocol with GPT-4o as the judge. Examples from 2026:
| Paper | Date (arXiv) | Reported PhyGenBench result |
|---|---|---|
| Chain of Event-Centric Causal Thought (CoECT) | March 2026 | Average 0.66 on a CogVideoX-5B base, against 0.61 for the previous best method it lists (PhysHPO)[10] |
| NEWTON: Agentic Planning for Physically Grounded Video Generation | May 2026 | LTX-Video-2B raised from 0.510 to 0.560 without retraining the planner on PhyGenBench; Wan2.2-TI2V-5B listed at 0.544[12] |
| PhysRAG | June 2026 | Wan 2.2 5B raised from 0.54 to 0.58, judged by GPT-4o under the official protocol[11] |
| VideoCoCo (Code-as-CoT) | July 2026 | OmniWeaving raised from 0.475 to 0.558, judged by GPT-4o under the official protocol[13] |
Benchmark papers also cite it as a reference point. The ICLR 2026 paper PhyWorldBench compares its 1,050 prompts with PhyGenBench's 160, and PhyGround (May 2026) drew part of its 13-law taxonomy from PhyGenBench's categories.[14][15]
NVIDIA Physis-Lang
NVIDIA's Physis-Lang paper, announced on 29 September 2026 and written by researchers from NVIDIA, MIT and the University of Oxford, uses PhyGenBench as one of its four physics benchmarks, alongside VideoPhy-2, PhyGround and Physics-IQ Verified.[6][7] It changes the judging protocol. The authors "replace the original offline VLM evaluators of VideoPhy-2 and PhyGenBench with GPT-5.5", saying it gives more discriminative assessments.[6] The appendix says they compute "the video-level score with the official scripts" but with GPT-5.5 as the evaluator: each video gets a single 0-3 grade, from completely fantastical to strongly realistic, averaged over all 160 prompts and rescaled to 100.[6] That is the same four-level scale PhyGenEval uses in its overall naturalness stage. The project page likewise notes that "PhyGenBench and VideoPhy-2 use GPT-5.5 evaluation".[7] Physis-Lang's PhyGenBench numbers therefore cannot be compared directly with scores from the original PhyGenEval ensemble or from papers using the GPT-4o protocol.
Physis-Lang's PhyGenBench table (GPT-5.5 judge, 100-point scale):[6]
| Model | Force | Light | Heat | Material | Overall |
|---|---|---|---|---|---|
| CogVideoX1.5-5B | 38.33 | 57.33 | 55.56 | 43.33 | 48.75 |
| Wan2.2-TI2V-5B | 40.83 | 60.67 | 43.33 | 40.83 | 47.50 |
| HunyuanVideo-1.5 | 47.50 | 56.67 | 44.44 | 41.67 | 48.33 |
| Veo 3.1 | 65.00 | 70.67 | 57.78 | 65.83 | 65.63 |
| PhysVid | 48.33 | 62.00 | 33.33 | 40.83 | 47.92 |
| PhyGDPO | 43.33 | 62.67 | 43.33 | 41.67 | 48.96 |
| Self-Refinement | 46.67 | 59.33 | 44.44 | 42.50 | 49.17 |
| Cosmos3-Nano | 64.17 | 64.00 | 57.78 | 59.17 | 61.67 |
| Physis-Lang (Cosmos3-Nano) | 67.50 | 72.67 | 75.56 | 69.17 | 71.04 |
On other backbones, the method raised PhyGenBench scores from 56.67 to 65.83 on Wan2.1-14B, from 48.96 to 52.50 on Cosmos3-Edge-4B and from 66.04 to 70.21 on Cosmos3-Super-64B.[6]
On 29 September 2026 NVIDIA's @NVIDIAAI account said that "adding physics reasoning to the prompt alone improved NVIDIA Cosmos 3's PhyGenBench score by 5.62 points without retraining."[8] In the paper, that figure matches Table 7: pretrained Cosmos3-Nano rose from 61.67 with the base caption to 67.29 with the generated physics reasoning plus a scene-specific physics negative prompt. The physics reasoning alone, without the negative prompt, raised the score to 63.33, a gain of 1.66 points, which the paper calls a zero-shot gain.[6] A manually written physics caption scored 61.86.[6]
Cosmos 3
NVIDIA's own Cosmos 3 technical report (arXiv 2606.02800, version 4 of 23 June 2026) evaluates video generation with PAIBench-G, RBench and Physics-IQ plus human evaluation, and uses VideoPhy2 annotations to train Cosmos 3's ability to judge physical plausibility. The PhyGenBench scores for NVIDIA Cosmos 3 models cited above come from the Physis-Lang paper, not from the Cosmos 3 report.[6][9]
Criticism and limitations
- Size and one-law-per-prompt design. PhyGround's authors write that PhyGenBench "curates only 160 manually crafted prompts and assumes each video corresponds to only one physical law", and they excluded its naturalness score from their own protocol because a holistic score gives no information about which law a model violated.[15]
- Closed-model judges. PhyGround also notes that PhyGenEval relies on closed-source models such as GPT-4o, "raising reproducibility concerns because model versions, weights, and API behavior change outside benchmark users' control."[15] The two different GPT-4o snapshots named in the README and the ICML paper, and Physis-Lang's switch to GPT-5.5, show how the judge can change between reports.[2][3][6]
- Fixed event order. The PhyWorldBench authors say PhyGenEval "assumes a fixed event order", which many of their own scenarios do not have, so they did not use it.[14]
- Domain axes and VLM judges. A June 2026 study of VLM judges (JudgeFit, "Each Judge Its Own Yardstick") found that under PhyGenBench's fixed four-domain schema, scores from six VLM judges shown in its first figure concentrated on the mechanics axis while the other three were "nearly inactive", giving little separation between models.[16]
- Scores that differ between papers. Results depend on the judge and pipeline, so numbers for the same base model vary across papers. The original paper scored CogVideoX-5B at 0.45; the NEWTON paper, with its own VLM-judged run, lists CogVideoX-5B at 0.354.[2][12]
Related benchmarks
PhyGenBench sits among several physics benchmarks for video generation. VideoPhy and VideoPhy-2 rate semantic adherence and physical commonsense with trained judges. Physics-IQ compares generated continuations with real recorded videos instead of using a VLM judge.[9] PhyWorldBench, PhyGround and PhysVidBench add more prompts, finer per-law scoring or tool-use scenarios.[14][15][17]
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33 ^34 ^35 ^36 ^37 ^38 ^39Fanqing Meng, Jiaqi Liao, Xinyu Tan, Wenqi Shao, Quanfeng Lu, Kaipeng Zhang, Yu Cheng, Dianqi Li, Yu Qiao, Ping Luo. "Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation". arXiv:2410.05363 (technical report), 7 October 2024. arxiv.org/...2410.05363
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16Fanqing Meng, Jiaqi Liao, Xinyu Tan, Quanfeng Lu, Wenqi Shao, Kaipeng Zhang, Yu Cheng, Dianqi Li, Ping Luo. "Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation". Proceedings of the 42nd International Conference on Machine Learning, PMLR 267:43781-43806, 2025. proceedings.mlr.press/...meng25c
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13OpenGVLab. "PhyGenBench" GitHub repository (README). Retrieved 30 September 2026. github.com/...PhyGenBench
- ^1 ^2"PhyGenBench: Physics Generation Benchmark". Project page. Retrieved 30 September 2026. phygenbench123.github.io
- ^ICML 2025. "Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation". Poster page. icml.cc/...44642
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10Liming Lu, Xianzheng Ma, Wenkun He, Guanqi Zhan, et al. "Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model". NVIDIA, MIT, University of Oxford, September 2026 (paper PDF). physis-intelligence.github.io/...Physis-Lang.pdf
- ^1 ^2Physis-Lang project page. NVIDIA, MIT, University of Oxford. Retrieved 30 September 2026. physis-intelligence.github.io/physis-lang-web
- ^NVIDIA AI (@NVIDIAAI). Post announcing Physis-Lang. X, 29 September 2026. x.com/...2105051552029556942
- ^1 ^2NVIDIA. "Cosmos 3: Omnimodal World Models for Physical AI". arXiv:2606.02800v4, 23 June 2026. arxiv.org/...2606.02800
- ^Zixuan Wang, Yixin Hu, Haolan Wang, Feng Chen, Yan Liu, Wen Li, Yinjie Lei. "Chain of Event-Centric Causal Thought for Physically Plausible Video Generation". arXiv:2603.09094, March 2026. arxiv.org/...2603.09094
- ^Kexu Cheng, Zicheng Liu, Mingju Gao, Chunhe Song, Hao Tang. "PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation". arXiv:2606.26916, 25 June 2026. arxiv.org/...2606.26916
- ^1 ^2Yuxiang Feng, Juncheng Wang, Chao Xu, et al. "NEWTON: Agentic Planning for Physically Grounded Video Generation". arXiv:2605.18396, May 2026. arxiv.org/...2605.18396
- ^Haodong Li, Tianfei Ren, Xiaoxiao Ma, et al. "VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System". arXiv:2607.27380, July 2026. arxiv.org/...2607.27380
- ^1 ^2 ^3Jing Gu, Xian Liu, Yu Zeng, et al. "PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models". ICLR 2026; arXiv:2507.13428v3. arxiv.org/...2507.13428
- ^1 ^2 ^3 ^4Juyi Lin, Arash Akbari, Yumei He, et al. "PhyGround: Benchmarking Physical Reasoning in Generative World Models". arXiv:2605.10806, 11 May 2026. arxiv.org/...2605.10806
- ^Yu Cao, Ziquan Liu, Zhensong Zhang, Jiankang Deng, Shaogang Gong, Jifei Song. "Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation". arXiv:2606.22918, 22 June 2026. arxiv.org/...2606.22918
- ^Enes Sanli, Baris Sarper Tezcan, Erkut Erdem, Aykut Erdem. "PhysVidBench: Language-Grounded Evaluation of Physical Commonsense in Text-to-Video Models". arXiv:2507.15824v2. arxiv.org/...2507.15824
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 3,661 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent verification (xg15 V2, 30 Sep 2026): checked vs arXiv and ICML PDFs, PMLR, GitHub README, 2026 adoption papers; 1 minor defect fixed in v2
Cite this page: AI Wiki. "PhyGenBench." aiwiki.ai, updated 30 Sept 2026, fact-checked 30 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/phygenbench