Megadose AI progress, ranked and analyzed.

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

· HF Daily Papers ·
The benchmark says top video models still fail at physics as reasoning, not just rendering.

Apple-PI tests video generators against classical mechanics through 400 annotated videos and a three-stage protocol: perception, formulation, and deduction. The authors treat the generated video as a visible reasoning trace, then score it with both MLLM judgments and physics-grounded measures. Across 11 models, the best video model scored 0.473, with failures clustering along the perception-to-deduction pipeline. HF Daily Papers' note

score 5

Categories: Research