Megadose AI progress, ranked and analyzed.

Can We Trust Item Response Theory for AI Evaluation?

· ArXiv · AI/CL/LG ·
IRT can mislead AI benchmark rankings when the model pool is small or oddly distributed.

The paper tests IRT methods under conditions closer to LLM benchmarking than human testing: fewer systems, many more items, and nonnormal capability spreads. Across 18,000 simulations, classical estimators sometimes become impractical at benchmark scale. Faster estimators can still give unreliable item-level and ranking conclusions. The authors frame IRT as usable only with enough sample size and diagnostics, not as a default trust layer for benchmark claims. ArXiv · AI/CL/LG's note

score 4

Categories: Research