NVIDIA 联合 MIT 与牛津发布 Physis-Lang,助力 Cosmos 3 在物理基准上超越 Veo 3.1
NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks
NVIDIA、MIT 和牛津团队发布 Physis-Lang,一个自进化的物理语言框架,通过在视频描述中加入 physics_reasoning 字段和 physics_negative_prompt,把物理语言用于数据筛选、训练和推理。
Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford argues the fix can come from language itself, not from extra visual, latent or numerical signals.
Their framework, Physis-Lang, treats physical language as a shared, optimizable representation. The same text drives data curation, model training and inference. On the public Physics-IQ Verified leaderboard snapshot dated September 29, 2026, Physis-Lang on Cosmos3-Super ranks first at 48.2 ± 1.4. The Cosmos3-Nano version ranks second at 43.3 ± 1.5.
A video world model can make a convincing clip and still get the physics wrong.
— NVIDIA AI (@NVIDIAAI) September 29, 2026
Our researchers just released Physis-Lang, an open self-evolving framework that adds physics reasoning to video captions. The captions explain why and how a scene unfolds. We use them to fine-tune… pic.twitter.com/agIHuIB6N2
What Problem Does Physis-Lang Solve?
Conventional captions describe what happens, not why. ‘Butter melts as the temperature rises’ says nothing about heat transfer or gravity. Physis-Lang adds a physics_reasoning field to each base caption. It spells out entities, causes, interactions, governing principles, temporal evolution and effects.
The pipeline also writes a scene-specific physics_negative_prompt. This text describes likely implausible outcomes, such as a stone floating on water. It acts as negative conditioning at inference time.
How Does the Self-Evolving Caption Loop Work?
The loop keeps the captioner frozen and evolves only its instruction. A GPT-5.5 captioner writes captions for a fixed 20-video development set with 273 human-verified assertions. Gemini-3.1-Pro acts as a physics-aware critic. An evolution agent reads the scores and claim-level failures, then rewrites the prompt.
The critic scores 2 dimensions:
- Precision: the caption is split into atomic claims, and each claim is checked against the video.
- Recall: each human-curated physical assertion must be explicitly stated or entailed by the caption.
Every revised prompt is validated on PhysCapBench, a new benchmark of 246 videos and 3,794 human-verified assertions. Caption F1 rose from 78.64 at iteration 1 to 87.82 at iteration 9. The path was not smooth. Iteration 2 made captions overly cautious and dropped F1 to 76.28. Iteration 9 required every visible causal step and raised frame sampling from 2 fps to 4 fps.
How Does Language-Guided Data Curation Work?
A GPT-5.5 diagnosis agent maps generated-video failures to physics categories like rigid-body motion, collision and fluid dynamics. That deficiency profile is matched against physics tags on a large video gallery. Retrieval targets physical content, not visual appearance.
The final training set holds 183K videos: 71K filtered from WISA-80K plus 112K retrieved clips. Retrieval alone added 3.01 points on average across 3 benchmarks. On VideoPhy-2, chemical and thermal processes each gained 8.00 points.
How Does Physis-Lang Compare With Veo 3.1?
Fine-tuning uses LoRA on attention projections, with no architecture or objective change. Physis-Lang on Cosmos3-Nano versus Google's Veo 3.1:
- PhyGenBench: 71.04 vs 65.63
- Physics-IQ Verified: 43.41 vs 34.99
- PhyGround: 69.90 vs 69.24
- VideoPhy-2: 68.02 vs 68.87 on the full set, 62.36 vs 58.43 on the Hard split
Gains hold across backbones: +7.05 on Wan2.1-14B, +3.24 on Cosmos3-Edge-4B, +6.22 on Cosmos3-Nano-16B and +5.02 on Cosmos3-Super-64B. General quality held steady on VBench-I2V, where Cosmos3-Nano moved from 88.32 to 88.69.
Prompting alone also helps. Physics reasoning plus negative prompts lifted a frozen Cosmos3-Nano on PhyGenBench from 61.67 to 67.29.
Can It Run Without Commercial APIs?
The research team distilled the GPT pipeline into 2 Qwen3-VL-4B-Instruct models: PhysThinker-C for captioning and PhysThinker-U for prompt upsampling. On Wan2.1-14B, the commercial pipeline gave +7.05 at about $24.12K in API cost. Swapping in PhysThinker-C kept +6.76 at about $0.12K. A fully local setup cost $0 and still added +4.76.
Physis-Lang vs Closest Competitors
| Feature | Physis-Lang | PhiZero | PhyGDPO | Self-Refinement |
|---|---|---|---|---|
| Developer | NVIDIA, MIT, Oxford | CASIA (NLPR) | Meta (ECCV 2026) | Liu et al. |
| Core idea | Self-evolving natural-language physics captions and negative prompts | Learned discrete "physical language", reason-then-render | Groupwise DPO with VLM physics rewards | Multimodal chain-of-thought prompt refinement from VLM feedback |
| Where physics enters | Data curation, training captions and inference prompts | Qwen3-VL-4B reasoner feeding a diffusion decoder | Preference training on PhyVidGen-135K | Inference prompts only |
| Training needed | LoRA SFT (prompt-only mode also helps) | Yes | Yes (DPO) | No, training-free |
| Physics-IQ Verified | 43.41 | 40.91 | n/r | 27.20 |
| PhyGenBench | 71.04 | n/r | 48.96 | 49.17 |
| VideoPhy-2 (All / Hard) | 68.02 / 62.36 | n/r | 59.56 / 44.94 | 47.88 / 28.09 |
| PhyGround | 69.90 | 57.85 | n/r | 58.22 |
| Code / weights public | Paper only | "Coming soon" | "Released soon" | Paper |
Scores are from the Physis-Lang paper, Tables 1 to 4, run under one protocol per benchmark (PhyGenBench and VideoPhy-2 use a GPT-5.5 evaluator). Physis-Lang numbers use the Cosmos3-Nano backbone. n/r = not reported in that comparison. Release status checked September 30, 2026.
Key Takeaways
- Physis-Lang evolves physics captions with a critic-guided agent while the captioner stays frozen.
- PhysCapBench scores captions on 3,794 human-verified cause, law and effect assertions.
- Cosmos3-Nano with Physis-Lang beats Veo 3.1 on 3 of 4 benchmarks.
- Physics prompts alone lift a frozen model by 5.62 points on PhyGenBench.
- No code or weights are public yet; only the paper is released.
Check out the Paper, Project Page and GitHub Repo. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks appeared first on MarkTechPost.
来源:MarkTechPost(RSS) · marktechpost.com