Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations
The paper proposes stopping LLM eval samples once uncertainty is low, instead of spending the same budget on every item.
Toby D. Pilditch’s `optstop` framework treats evaluation as sequential measurement, using hierarchical Bayesian inference to decide where more trials are still useful. It supports binary, ordinal, and continuous outcomes while keeping all benchmark items eligible for sampling. In a 200-item, 10-epoch example, it cut 57% to 97% of planned trials across nine validation settings while preserving the full-run conclusions. ArXiv · AI/CL/LG's note
Toby D. Pilditch’s `optstop` framework treats evaluation as sequential measurement, using hierarchical Bayesian inference to decide where more trials are still useful. It supports binary, ordinal, and continuous outcomes while keeping all benchmark items eligible for sampling. In a 200-item, 10-epoch example, it cut 57% to 97% of planned trials across nine validation settings while preserving the full-run conclusions. ArXiv · AI/CL/LG's note
score 5