Megadose AI progress, ranked and analyzed.

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

· HF Daily Papers ·
VeriPhy judges generated videos by tracing specific physical failures to typed evidence, not by issuing a single quality score.

The paper describes a system that turns a prompt into physical obligations before seeing the frames, then checks only declared measurements such as tracking, counting, depth, OCR, and audio events. Its verdicts resolve to plausible, implausible, or abstain, with provenance attached to each evidence record. The authors evaluate it on a 1,500-clip annotated corpus and say that, on a 149-clip core set, VeriPhy accounts for 228 flaw records versus 164 for a prior question-decomposition evaluator. A monolithic prompt reaches 222, but the paper’s claimed distinction is auditability and the ability to feed traced critic verdicts back into generation. HF Daily Papers' note

score 5

Categories: Research