Megadose AI progress, ranked and analyzed.

Item Response Theory for AI Safety

· ArXiv · AI/CL/LG ·
The paper argues that IRT can make AI safety benchmarks cheaper to run and harder to game.

The authors fit Item Response Theory models across eight safety benchmarks and 192 language models. They find three main factors behind benchmark variance: refusal strictness, truthfulness, and contextual harm. The paper says psychometrically selected questions can approximate full benchmark scores with less error than random subsets, with about ten adaptive items enough for several benchmarks. It also claims IRT can help audit individual models for naive sandbagging and API model changes. ArXiv · AI/CL/LG's note

score 4

Categories: Research