Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
The benchmark says top video models still fail at physics as reasoning, not just rendering.
Apple-PI tests video generators against classical mechanics through 400 annotated videos and a three-stage protocol: perception, formulation, and deduction. The authors treat the generated video as a visible reasoning trace, then score it with both MLLM judgments and physics-grounded measures. Across 11 models, the best video model scored 0.473, with failures clustering along the perception-to-deduction pipeline. HF Daily Papers' note
Apple-PI tests video generators against classical mechanics through 400 annotated videos and a three-stage protocol: perception, formulation, and deduction. The authors treat the generated video as a visible reasoning trace, then score it with both MLLM judgments and physics-grounded measures. Across 11 models, the best video model scored 0.473, with failures clustering along the perception-to-deduction pipeline. HF Daily Papers' note
score 5