Can We Trust Item Response Theory for AI Evaluation?
IRT can mislead AI benchmark rankings when the model pool is small or oddly distributed.
The paper tests IRT methods under conditions closer to LLM benchmarking than human testing: fewer systems, many more items, and nonnormal capability spreads. Across 18,000 simulations, classical estimators sometimes become impractical at benchmark scale. Faster estimators can still give unreliable item-level and ranking conclusions. The authors frame IRT as usable only with enough sample size and diagnostics, not as a default trust layer for benchmark claims. ArXiv · AI/CL/LG's note
The paper tests IRT methods under conditions closer to LLM benchmarking than human testing: fewer systems, many more items, and nonnormal capability spreads. Across 18,000 simulations, classical estimators sometimes become impractical at benchmark scale. Faster estimators can still give unreliable item-level and ranking conclusions. The authors frame IRT as usable only with enough sample size and diagnostics, not as a default trust layer for benchmark claims. ArXiv · AI/CL/LG's note
score 4