Megadose Built for builders and researchers.

On the estimation and validity of AI time horizons---a statistical look at the METR plot

· ArXiv · AI/CL/LG ·
The paper argues METR-style time horizons are not evenly meaningful across the scale.

Nguyen and Fithian recompute 50% AI time horizons on 228 tasks and 26 AIs using splines and item-response theory. Their fit finds human task time maps to AI difficulty almost flat from about 2 to 30 minutes, but closer to linear outside that range. That means a jump from 3 to 30 minutes is statistically easier than a jump from 30 minutes to 5 hours, even though both are 10x increases. They recommend reading time-horizon estimates alongside diagnostic plots as benchmarks add longer tasks. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research