Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model
The paper argues that language can carry the physical rules a video model needs, then uses that language to improve generation.
Physis-Lang describes scenes in terms of entities, causes, interactions, principles, time changes, and effects. The authors add PhysCapBench to score those captions as atomic physical assertions, then run an agentic loop that refines the captioning instructions from the errors. They also turn model failures into text and retrieve videos covering the missing physical processes. On four physical-video benchmarks using Wan and Cosmos backbones, the authors report better physical plausibility, including Physis-Lang-enhanced Cosmos3-Nano models beating Veo 3.1.
HF Daily Papers' note
Physis-Lang describes scenes in terms of entities, causes, interactions, principles, time changes, and effects. The authors add PhysCapBench to score those captions as atomic physical assertions, then run an agentic loop that refines the captioning instructions from the errors. They also turn model failures into text and retrieve videos covering the missing physical processes. On four physical-video benchmarks using Wan and Cosmos backbones, the authors report better physical plausibility, including Physis-Lang-enhanced Cosmos3-Nano models beating Veo 3.1.
HF Daily Papers' note
score 5