A Statistical Audit of Physical AI Benchmark Redundancy
Two substitute benchmark pairs were redundant enough to materially change model rankings.
The paper builds a 51-model, 12-benchmark matrix for physical AI and measures overlap between benchmark scores. When the authors collapse the two substitute pairs into single columns, 22 of 51 models move by at least three ranking places under an equal-weight average. A greedy selection method finds four benchmarks that retain 78.5% of the utility of all 12, then uses that subset for a Bradley-Terry ranking. ArXiv · AI/CL/LG's note
The paper builds a 51-model, 12-benchmark matrix for physical AI and measures overlap between benchmark scores. When the authors collapse the two substitute pairs into single columns, 22 of 51 models move by at least three ranking places under an equal-weight average. A greedy selection method finds four benchmarks that retain 78.5% of the utility of all 12, then uses that subset for a Bradley-Terry ranking. ArXiv · AI/CL/LG's note
score 5