Physis-Lang
Physis-Lang is a research framework from NVIDIA, the Massachusetts Institute of Technology and the University of Oxford that uses structured natural language as a representation of physical processes for video world models. Its paper, "Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model", describes a pipeline in which physics-rich captions (stating cause, governing physical law and effect) are refined by an agentic loop, used to find missing physics in training data, used to fine-tune video generators, and used again at inference time to expand prompts and supply scene-specific negative descriptions.[1][2] NVIDIA's @NVIDIAAI account announced the work on 29 September 2026, saying that adding physics reasoning to the prompt alone raised NVIDIA Cosmos 3's PhyGenBench score by 5.62 points without retraining.[3] Two fine-tuned Cosmos 3 models built with the method held first and second place on the public Physics-IQ Verified image-to-video leaderboard as of 30 September 2026, at 48.2% and 43.3%; the results were self-reported by the authors.[1][5][7]
The authors argue against a common assumption in physics-aware video generation, namely that language is too weak to carry physical knowledge and must be supplemented with visual, latent, numerical or planning-based signals. They propose instead that language "when explicitly structured and self-evolving" can serve as the representation that carries physical knowledge across data curation, training and inference.[2]
Overview
| Attribute | Detail |
|---|---|
| Full title | Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model[2] |
| Institutions | NVIDIA, MIT, University of Oxford[2] |
| Corresponding authors | Guanqi Zhan and Han Cai (NVIDIA)[2] |
| Public announcement | 29 September 2026 (NVIDIA AI post on X)[3] |
| Paper | PDF on the project page and in the GitHub repository; copyright line "(c) 2026 NVIDIA. All rights reserved."[2][4] |
| Backbones tested | Wan2.1-14B; Cosmos3-Edge (4B), Cosmos3-Nano (16B), Cosmos3-Super (64B)[2] |
| Training data | About 183K real-world videos (71K filtered from WISA-80K plus 112K retrieved)[2] |
| New benchmark | PhysCapBench: 246 videos, 3,794 human-verified assertions[2] |
| Evaluation benchmarks | PhyGenBench, VideoPhy-2, PhyGround, Physics-IQ Verified[2] |
| Code status (30 Sep 2026) | Repository holds the paper and project references only; no implementation or evaluation code[4] |
Authors and affiliations
The author list is identical on the paper and the project page. Four authors are marked as equal contributors, and two are marked as corresponding authors.[1][2]
| Author | Affiliation | Role noted in paper |
|---|---|---|
| Liming Lu | MIT | Equal contribution |
| Xianzheng Ma | University of Oxford | Equal contribution |
| Wenkun He | MIT | Equal contribution |
| Guanqi Zhan | NVIDIA | Equal contribution; corresponding author |
| Yilin Zhao | NVIDIA | |
| Junyu Chen | NVIDIA | |
| Mengyao Xu | NVIDIA | |
| Jiaojiao Fan | NVIDIA | |
| Wenhang Ge | NVIDIA | |
| Yuchao Gu | NVIDIA | |
| Yunze Liu | NVIDIA | |
| Boyi Li | NVIDIA | |
| Zhen Dong | NVIDIA | |
| Victor Prisacariu | University of Oxford | |
| Ming-Yu Liu | NVIDIA | |
| Song Han | NVIDIA | |
| Han Cai | NVIDIA | Corresponding author |
The paper gives the corresponding authors' contact addresses as gzhan@nvidia.com and hcai@nvidia.com.[2] The suggested BibTeX entry (key lu2026physislang) lists the institution as NVIDIA.[1]
Motivation
The paper opens from a familiar failure of video generators: they produce clips that look realistic but break basic physics, with objects deforming or vanishing, collisions producing implausible outcomes, fluids moving unnaturally and causal events happening in the wrong order. The authors attribute this partly to conventional video-text training, which optimizes visual fidelity and semantic alignment without explicitly representing the physical knowledge behind the observed dynamics.[2]
The related-work section groups prior physics-aware video generation methods into three lines: work that draws physical priors from pretrained vision-language or foundation models (distilling dynamics into auxiliary branches, retrieving reference motions, critiquing outputs, or planning over intermediate representations); work that optimizes physical plausibility with reinforcement learning or preference optimization; and work that injects simulator- or graphics-derived structure such as equations, trajectories or simulation code. Some works also curate physics-annotated datasets. According to the authors, most of these approaches treat language as a way to specify semantic content and route physical detail through other channels.[2]
Physis-Lang tests the opposite hypothesis: that language's limited role in current video models comes from language not yet being fully exploited as a representation of physical processes, not from an intrinsic limit of language.[2] The paper's related-work section goes further and states that the authors "prove that language alone is very powerful to lead to physically plausible video generation if we utilize it properly", and claims to be the first to apply agentic self-evolution to physics reasoning for video generation.[2] Both are the authors' own characterizations.
The physical-language representation
A conventional caption describes what happens. A Physis-Lang caption also explains why. The paper's running example contrasts the conventional caption "Butter melts as the temperature rises" with a three-part physical caption:[1][2]
| Part | Example text |
|---|---|
| Cause | Rising temperature supplies heat to the butter. |
| Physical law | Heat drives melting; gravity makes the butter slump. |
| Effect | The solid shrinks as a shallow liquid pool expands. |
In the implementation, the base caption is augmented with a natural-language field called physics_reasoning that describes physical details, dynamics and causal relations. A second field, physics_negative_prompt, describes physically implausible ways the scene is likely to evolve, and is used as negative conditioning during generation.[2] More broadly, the abstract says the language describes a process's "relevant entities, causes, interactions, governing principles, temporal evolution, and effects".[2]
Framework
The project page presents the method as a single representation used in three connected stages.[1]
Stage A: self-evolving captioning guidelines
The first stage improves the instruction (the "guidelines") given to a captioning model, not the captioner itself. The captioner's weights are never updated.[2][4]
- A fixed captioner, instantiated with GPT-5.5, receives temporally ordered frames and generates
physics_reasoningandphysics_negative_promptin one call. Onlyphysics_reasoningis scored.[2] - A physics-aware critic scores the captions on a fixed development set of 20 videos, randomly sampled from the training set and carrying 273 human-verified atomic physics assertions. Within the loop, Gemini 3.1 Pro measures assertion-level recall and video-grounded claim precision.[2]
- An evolution agent reads the aggregate scores and the claim-level failure rationales, looking for omissions, unsupported claims and vague causal relations, and rewrites the captioning prompt. The paper writes the cycle as P_t -> C_t -> R_t -> P_(t+1) (prompt, captions, report, next prompt).[2]
- After each iteration the new prompt is scored on PhysCapBench, which serves as the held-out validation set and stopping criterion. The loop continues while benchmark performance improves and stops when the gain saturates; the best-scoring prompt becomes the final instruction.[2]
An appendix annotates the prompt changes at several iterations. At iteration 2 the agent added stricter grounding and causal-chain rules, which made captions "too conservative" and caused them to miss useful physical details. Iteration 3 added detailed checklists for tracking objects, contacts and event stages. Iteration 4 simplified the rigid checklists and focused on the main physical processes. Iteration 8 required stronger visual evidence for every event. Iteration 9, the best-scoring version, required every visible causal step to be stated explicitly and raised the frame sampling rate from 2 fps to 4 fps.[2]
Stage B: language-guided retrieval of missing physics
The second stage uses language to decide what additional training data to collect. A diagnosis agent built on GPT-5.5 identifies physical failures in the target model's generated videos and assigns each failure to physics categories such as rigid-body motion, collision and fluid dynamics. Aggregating these gives a category-level "deficiency profile".[2]
Videos in a large gallery carry physics-domain tags derived from their captions. The system matches the deficient categories against those tags, so retrieval is driven by physical content instead of visual similarity. The authors argue this lets the data engine collect visually diverse examples of the same missing physics. Retrieved videos are re-captioned with the evolved guidelines and added to the training set.[2]
Stage C: training and generation
The evolved language is reused twice more:[1][2]
- Training. The final captioner re-captions the training videos. Each training caption combines the base visual description with the generated physics reasoning.
- Inference. An "upsampler" built from the same final caption prompt expands a short benchmark prompt, and an input image when the task requires one, into the same structured
physics_reasoningandphysics_negative_promptfields. The positive reasoning gives the generator explicit physical guidance, and the negative prompt is used to suppress physically implausible evolutions.
Because physics-rich prompts are long, the authors added a sliding-window scheme for Wan2.1, whose UMT5-XXL text encoder has a 512-token limit. Over-length prompts are split into overlapping 512-token windows with a 192-token overlap, and embeddings in the overlapping region are averaged.[2]
PhysThinker
The captioner and upsampler are first instantiated with proprietary models, which the paper says are costly at the scale of re-captioning a training corpus. To reduce cost, the authors distilled them into two 4B vision-language models called PhysThinker: PhysThinker-C for physics-aware video captioning and PhysThinker-U for prompt upsampling. Both are Qwen3-VL-4B-Instruct models fine-tuned with LoRA on the language model and full fine-tuning of the vision stack. The captioner was trained on 128 A100 GPUs and the upsampler on 32 A100 GPUs.[2]
PhysCapBench
PhysCapBench is a benchmark the authors built to measure how much physical detail a video caption carries, and how much of it is correct.[1][2]
| Property | Value |
|---|---|
| Videos | 246 physics-rich videos, collected from Physics-IQ and YouTube |
| Assertions | 3,794 human-verified atomic physical assertions |
| Density | 15.4 assertions per video on average (15.42 in the appendix) |
| Assertion types | Causes, relevant physical laws and effects |
| Construction | AI annotators propose atomic assertions; humans remove incorrect or ambiguous ones and add missing ones |
| Evaluator | Gemini-3-Flash, temperature 0.1, batches of at most 30 claims |
Scoring is inspired by the caption evaluation protocol used for Cosmos 3, with precision and recall specialized to physical content:[2]
- Precision asks whether each claim is supported. The critic decomposes the caption into atomic, independently verifiable claims, checks each against the full video, and labels it correct, incorrect or uncertain. Precision = N_correct / (N_correct + N_incorrect), micro-averaged, so uncertain claims do not count either way.
- Recall asks whether the physical meaning is covered. The critic checks each human-curated assertion against the caption; an assertion passes only when the caption explicitly states or unambiguously entails its complete physical meaning. Recall = N_pass / N_ground-truth.
- F1 is the harmonic mean of the two.
A worked example in the paper uses a cup breaking on impact, with ground-truth assertions for the cause (the cup collides with a rigid obstacle), the law (the collision impulse changes the momentum of each part of the cup) and the effect (the cup body ruptures at the impact site).[2]
Experimental setup
Data and backbones
The training set contains about 183K real-world video-caption pairs: 71K samples kept from WISA-80K after filtering for motion and aesthetic quality and removing multi-shot or metadata-incomplete videos, plus 112K videos selected by the deficiency-guided retrieval pipeline.[2] The authors fine-tuned existing backbones without changing their architectures or training objectives:[2]
| Backbone | Parameters | Fine-tuning details reported |
|---|---|---|
| Wan2.1 T2V-14B and I2V-14B (tuned separately) | 14B | LoRA rank 256 on attention projections, patch embedder and output modules fully tuned; 16 NVIDIA H100 GPUs; checkpoint at iteration 8,000 |
| Cosmos3-Edge | 4B | LoRA rank 16; 480-resolution buckets; 32 GPUs on eight nodes; checkpoint at iteration 2,000 |
| Cosmos3-Nano | 16B | LoRA rank 16 on attention projections; 256 x 256 video; 16 GPUs on two nodes; checkpoint at iteration 1,340 |
| Cosmos3-Super | 64B | LoRA rank 16 on the generation expert's attention layers; 256-resolution buckets; 16 GPUs on four nodes; checkpoint at iteration 1000 |
The Cosmos 3 models support both text-to-video and image-to-video in one backbone. For Cosmos3-Nano and Super, training mixed 70% text-to-video, 20% image-to-video and 10% video-to-video conditioning.[2]
Benchmarks and evaluator
The paper evaluates on four existing physics benchmarks:[2]
| Benchmark | Task | Size and scoring as described in the paper |
|---|---|---|
| PhyGenBench | Text-to-video | 160 prompts in four categories (force, light, heat, material); each video graded 0-3; official scripts with the evaluator replaced by GPT-5.5 |
| VideoPhy-2 | Text-to-video | 591 test samples, 178 in a Hard subset; semantic adherence (SA) and physical commonsense (PC) graded 0-5; joint score is the fraction with SA >= 4 and PC >= 4; evaluator replaced by GPT-5.5 |
| PhyGround | Image-to-video | 250 samples; semantic adherence (General) and physical temporal validity (Physics), each scored 1-5 |
| Physics-IQ Verified | Image-to-video | 66 physical scenarios, three perspectives each (198 samples); spatial, spatiotemporal, weighted-spatial and MSE metrics against ground-truth videos |
All scores are rescaled to a 100-point scale. The replacement of the original offline VLM judges of VideoPhy-2 and PhyGenBench with GPT-5.5 is the most consequential protocol choice. The authors justify it in an appendix: the original VideoPhy-2 auto-evaluator gave nearly identical scores to very different models (22.50 to 23.69 across CogVideoX1.5-5B, Wan2.1-T2V-14B, Cosmos3-Nano and Veo 3.1), while GPT-5.5 separated them (49.41 to 68.87), and the authors say GPT-5.5's judgments "align with human judgment in most cases".[2] The paper states that all models were evaluated with the same protocol within each benchmark.[2]
Baselines were CogVideoX1.5-5B, Wan2.2-TI2V-5B, HunyuanVideo-1.5 and Google's closed Veo 3.1, plus physics-focused methods PhysVid, PhyGDPO, Self-Refinement, Kandinsky-WM and PhiZero.[2]
Results
All results below are self-reported by the authors.
Caption self-evolution
On PhysCapBench, F1 rose from 78.64 at iteration 1 to 87.82 at iteration 9 (+9.18 points).[1][2] The curve was not monotonic; the project page notes that caption F1 "is non-monotonic" and that iteration 9 was selected.[1] To check whether better captions translate into better video, the authors kept pretrained Cosmos3-Nano fixed and changed only the inference captions. Captions from iterations 1, 4, 8 and 9 gave PhyGenBench scores of 64.17, 65.63, 65.83 and 67.29, a gain of 3.12 points from iteration 1 to iteration 9.[1][2]
Main comparison (Cosmos3-Nano)
Tables 1 to 4 of the paper compare the Physis-Lang fine-tuned Cosmos3-Nano with the pretrained model and with Veo 3.1:[1][2]
| Model | PhyGenBench | Physics-IQ Verified | VideoPhy-2 All | VideoPhy-2 Hard | PhyGround |
|---|---|---|---|---|---|
| Cosmos3-Nano (pretrained) | 61.67 | 40.23 | 60.41 | 48.31 | 65.18 |
| Veo 3.1 | 65.63 | 34.99 | 68.87 | 58.43 | 69.24 |
| Physis-Lang (Cosmos3-Nano) | 71.04 | 43.41 | 68.02 | 62.36 | 69.90 |
The paper claims three things from these tables: consistent improvement over the base model on all four benchmarks, the strongest results among open-source models and methods tested, and better scores than Veo 3.1 on Physics-IQ Verified, PhyGround and PhyGenBench. On VideoPhy-2 the fine-tuned model trails Veo 3.1 by 0.85 points on the full set (68.02 vs 68.87) and leads it by 3.93 points on the Hard subset (62.36 vs 58.43).[2]
The per-category PhyGenBench breakdown (Table 3) shows the largest gain in the heat category:[2]
| Model | Force | Light | Heat | Material | Overall |
|---|---|---|---|---|---|
| Cosmos3-Nano | 64.17 | 64.00 | 57.78 | 59.17 | 61.67 |
| Veo 3.1 | 65.00 | 70.67 | 57.78 | 65.83 | 65.63 |
| Physis-Lang (Cosmos3-Nano) | 67.50 | 72.67 | 75.56 | 69.17 | 71.04 |
On Physics-IQ Verified (Table 1), the strongest physics-focused baseline was PhiZero at 40.91, which the fine-tuned Nano (43.41) exceeded. On PhyGround (Table 2), the fine-tuned Nano scored 71.54 General and 68.26 Physics.[2]
Gains across backbones
Table 5 reports the improvement from Physis-Lang on each backbone:[2]
| Backbone | PhyGenBench | Physics-IQ Verified | VideoPhy-2 | PhyGround | Mean gain |
|---|---|---|---|---|---|
| Wan2.1-14B | 56.67 -> 65.83 | 27.87 -> 35.15 | 57.02 -> 65.65 | 61.52 -> 64.64 | +7.05 |
| Cosmos3-Edge-4B | 48.96 -> 52.50 | 32.80 -> 34.69 | 32.99 -> 40.61 | 66.66 -> 66.56 | +3.24 |
| Cosmos3-Nano-16B | 61.67 -> 71.04 | 40.23 -> 43.41 | 60.41 -> 68.02 | 65.18 -> 69.90 | +6.22 |
| Cosmos3-Super-64B | 66.04 -> 70.21 | 45.92 -> 50.00 | 60.91 -> 72.42 | 68.67 -> 68.98 | +5.02 |
The mean gain is positive for every backbone, but individual cells vary: Cosmos3-Edge's PhyGround score fell slightly (66.66 to 66.56) and Cosmos3-Super's PhyGround gain was 0.31 points. The project page notes that "individual benchmark gains can vary".[1][2] The gain did not grow with scale inside the Cosmos 3 family; Nano gained more than Super on average.[2]
Language-guided data curation
Adding the 112K retrieved videos to the WISA-derived set (both fine-tuned on Cosmos3-Nano with the full prompt) raised scores on three benchmarks (Table 6):[2]
| Benchmark | WISA only | WISA + retrieved | Gain |
|---|---|---|---|
| PhyGenBench | 68.12 | 71.04 | +2.92 |
| Physics-IQ Verified | 40.68 | 43.41 | +2.73 |
| VideoPhy-2 | 64.63 | 68.02 | +3.39 |
| Mean | 57.81 | 60.82 | +3.01 |
Broken down by physical category on VideoPhy-2 (Figure 5, joint-score gain in percentage points; sample counts in parentheses, and a sample can belong to several categories): cloth deformation (167) +7.19, fracture mechanics (94) +7.45, elasticity (81) +1.23, chemical processes (50) +8.00, rigid-body motion (33) +6.06, contact/collision (29) +3.45, soft-body motion (27) +7.41, thermal processes (25) +8.00.[1][2]
Prompt-only versus fine-tuned gains
Table 7 separates what the language does at inference time from what fine-tuning adds. All rows are PhyGenBench scores on Cosmos3-Nano.[2]
| ID | Fine-tuning data | Inference prompt | PhyGenBench |
|---|---|---|---|
| A | None (pretrained) | Base caption | 61.67 |
| B | None | Manually designed physics caption | 61.86 |
| C | None | Base + physics_reasoning | 63.33 |
| D | None | Base + physics_reasoning + physics_negative_prompt | 67.29 |
| E | WISA | Base + physics_reasoning | 65.62 |
| F | WISA | Base + physics_reasoning + physics_negative_prompt | 68.12 |
| G | WISA + retrieved | Base + physics_reasoning | 66.88 |
| H | WISA + retrieved | Base + physics_reasoning + physics_negative_prompt | 71.04 |
The 5.62-point figure in NVIDIA's announcement corresponds to row A to row D: 61.67 to 67.29 with the pretrained weights unchanged.[2][3] That setting uses both the evolved physics_reasoning and the scene-specific physics_negative_prompt. The positive reasoning alone (A to C) added 1.66 points, which the paper describes as a zero-shot gain, and the agent-evolved prompt beat a manually designed physics prompt (B, 61.86).[2] Row D's 67.29 is the same number as the iteration-9 point in the self-evolution experiment.[2]
With prompts held fixed, fine-tuning added 3.55 points (C to G) and 3.75 points (D to H). The negative prompt added 2.50 points on the WISA-only model (E to F) and 4.16 points on the model trained with retrieved data (G to H). The full pipeline (H, 71.04) is 9.37 points above the pretrained baseline (A).[2]
Cost and the open PhysThinker pipeline
Table 8 compares commercial and local models for captioning and upsampling on Wan2.1-14B. Cost covers API spending for re-captioning training data and upsampling benchmark prompts.[2]
| Captioner | Upsampler | Mean gain | Approximate cost (USD) |
|---|---|---|---|
| Commercial | Commercial | +7.05 | $24.12K |
| PhysThinker-C | Commercial | +6.76 | $0.12K |
| PhysThinker-C | PhysThinker-U | +4.76 | $0 |
Replacing the commercial models with the local 4B models kept most of the gain at a fraction of the external cost, but the fully local pipeline gave up about a third of the mean improvement.[2]
General video quality and attention analysis
To check whether physics-heavy captions hurt ordinary generation, the authors ran VBench image-to-video. Wan2.1-14B went from 86.86 to 87.51 and Cosmos3-Nano from 88.32 to 88.69 after Physis-Lang fine-tuning.[2] An attention-map analysis on Cosmos3-Nano reported that after fine-tuning, attention to physical keywords in the reasoning text rose; for example, the attention score for "HEAT TRANSFER" in one caption increased by 52.2% and for "BENDS" (light refraction) by 48.7%.[2] The paper also shows qualitative driving and robotics demos, which are illustrative examples, not benchmark results.[2]
Physics-IQ Verified leaderboard
Physics-IQ Verified is an audited version of the Physics-IQ benchmark; its public leaderboard is run by Anates Labs, which also publishes a page documenting the dataset fixes.[5] As of 30 September 2026, its image-to-video ranking, sorted by mean score, placed two Physis-Lang entries at the top, both listed as NVIDIA, open source, using an LLM, and dated 25 September 2026, the day the authors submitted them (the pull request adding them was merged on 28 September):[5][7]
| Rank | Model | Physics-IQ Verified score |
|---|---|---|
| 1 | Physis-Lang (Cosmos3 Super) | 48.2% +/- 1.4 |
| 2 | Physis-Lang (Cosmos3 Nano) | 43.3% +/- 1.5 |
| 3 | Cosmos3 Super | 42.7% +/- 0.8 |
| 4 | Seedance 2.5 | 42.4% +/- 0.8 |
| 5 | MiniMax H3 | 39.8% +/- 0.3 |
| 6 | Cosmos3 Nano | 37.3% +/- 0.9 |
The Physis-Lang Super entry also led all four submetric boards (spatial 59.9, spatiotemporal 41.6, weighted spatial 48.5, MSE 43.0).[5] The project page states that the Super entry improves on the previous leading score of 42.7 by 5.5 percentage points, and that the Nano entry surpasses the original Cosmos3-Super by mean score.[1]
The leaderboard numbers are not the same as the paper's own Physics-IQ Verified figures. The paper reports Cosmos3-Nano at 40.23 and its Physis-Lang version at 43.41, and Cosmos3-Super at 45.92 rising to 50.00, while the leaderboard lists 37.3 and 43.3 for Nano and 42.7 and 48.2 for Super.[2][5] The leaderboard submission used its own setup: four seeds per model at 1280x720 and 24 FPS, prompts made of base captions plus physical reasoning (upsampled with GPT-5.5 through an API), and the Cosmos default negative prompt in place of the scene-specific physics negative prompt.[7] The leaderboard's Veo entries are Veo 3.1 Lite (31.8%) and Veo 3.1 Fast (30.0%) (its other Google entries are Gemini Omni 1.1 Flash at 35.3% and Gemini Omni Flash at 33.4%); the paper's Veo 3.1 figure of 34.99 comes from the authors' own evaluation.[2][5]
Release and availability
The GitHub repository Physis-Intelligence/Physis-Lang was created on 25 September 2026 and last updated on 30 September 2026; its commit history through that date consists of adding and updating the README and the paper PDF.[4] Its README says: "This repository currently provides the paper and project references. It does not currently contain the model implementation or evaluation code." As of 30 September 2026 the repository had no license file.[4] NVIDIA's announcement described Physis-Lang as "an open self-evolving framework", and the Physics-IQ Verified leaderboard labels the two entries "Open source".[3][5] The Cosmos 3 backbones themselves are open models, but the fine-tuned Physis-Lang checkpoints had not been released as of 30 September 2026, and the leaderboard submission refers to "a privately available video model".[2][4][7] No PhysThinker weights or PhysCapBench data had been published in the repository at that date.[4]
Limitations and caveats
The paper has no separate limitations section. The following caveats come from its own tables and from the release materials:
- Changed evaluators. PhyGenBench and VideoPhy-2 scores use GPT-5.5 in place of the benchmarks' original judges, so they cannot be compared directly with scores published under the official evaluators.[2]
- Self-reported numbers. The benchmark comparisons, including those against Veo 3.1, were run by the authors. The Physics-IQ Verified leaderboard entries are not an independent check: the submitters added them through a GitHub pull request describing the results as "self-reported and submitted for verification", and the leaderboard lists different absolute numbers from the paper.[5][7]
- Uneven per-benchmark gains. Veo 3.1 still leads on the VideoPhy-2 full set, and Cosmos3-Edge's PhyGround score dropped slightly after fine-tuning.[2]
- Reliance on proprietary models. The captioner, upsampler and diagnosis agent were GPT-5.5, and critics used Gemini models; the local PhysThinker replacement recovers only part of the gain (+4.76 vs +7.05 on Wan2.1-14B).[2]
- The headline prompt-only number includes negative prompting. The 5.62-point PhyGenBench gain requires both the physics reasoning and the negative prompt; positive reasoning alone gave 1.66 points.[2]
- No code yet. At release the repository contained only the paper.[4]
A summary by the AlphaSignal newsroom noted that the full pipeline's results mix language refinement, data selection, training and prompting, and listed general simulation ability, long-horizon consistency and performance on unseen physical processes as "open evaluation questions".[6]
Reception
AlphaSignal ran a newsroom article on 29 September 2026 headlined "NVIDIA's Physis-Lang Teaches Video AI Real Physics, Beats Google's Veo 3.1", repeating the leaderboard figures and the 5.62-point prompt-only result, and describing the 5.62-point ablation as measuring "prompt expansion with fixed weights".[6]
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15Physis-Lang project page, "Physis-Lang | Self-Evolving Language for Video World Models". NVIDIA, MIT, University of Oxford. Retrieved 30 September 2026. physis-intelligence.github.io/physis-lang-web
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33 ^34 ^35 ^36 ^37 ^38 ^39 ^40 ^41 ^42 ^43 ^44 ^45 ^46 ^47 ^48 ^49 ^50 ^51 ^52 ^53 ^54 ^55 ^56 ^57 ^58 ^59 ^60 ^61 ^62 ^63 ^64 ^65 ^66 ^67 ^68Liming Lu, Xianzheng Ma, Wenkun He, Guanqi Zhan, et al. "Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model". NVIDIA, 2026 (paper PDF, version of 29 September 2026). physis-intelligence.github.io/...Physis-Lang.pdf
- ^1 ^2 ^3 ^4NVIDIA AI (@NVIDIAAI). Post announcing Physis-Lang. X, 29 September 2026. x.com/...2105051552029556942
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Physis-Intelligence/Physis-Lang. GitHub repository (README and commit history). Retrieved 30 September 2026. github.com/...Physis-Lang
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Anates Labs. "Physics-IQ Verified" leaderboard, I2V Verified Ranking. Retrieved 30 September 2026. physics-iq-verified.anates.ai
- ^1 ^2AlphaSignal Newsroom. "NVIDIA's Physis-Lang Teaches Video AI Real Physics, Beats Google's Veo 3.1". AlphaSignal, 29 September 2026. alphasignal.ai/...l-physics-beats-google-s-veo-3-1
- ^1 ^2 ^3 ^4 ^5"Add Physis-Lang (Cosmos3 Nano) and Physis-Lang (Cosmos3 Super) to Physics-IQ Verified". Pull request 83, google-deepmind/physics-IQ-benchmark, GitHub, opened 25 September 2026, merged 28 September 2026. github.com/...83
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
1 revision · v2 · 4,317 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent verification (xg15 V2, 30 Sep 2026): ~230 claims vs paper, project page, repo, PR #83, tweet; 3 material (self-reported leaderboard entries, Google entries, quote citation) + 6 minor defects fixed in v2
Cite this page: AI Wiki. "Physis-Lang." aiwiki.ai, updated 30 Sept 2026, fact-checked 30 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/physis_lang