Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
Model rankings changed when the token budget changed.
The paper tested four models across seven generation limits, from 64 to 4,096 tokens, on three reasoning benchmarks. It found that more room to answer did not always help: 3–19% of items became wrong at higher budgets even after controlling for truncation. Rankings reversed across budgets on every benchmark, and the authors argue evaluations should report results conditioned on inference budget. ArXiv · AI/CL/LG's note
The paper tested four models across seven generation limits, from 64 to 4,096 tokens, on three reasoning benchmarks. It found that more room to answer did not always help: 3–19% of items became wrong at higher budgets even after controlling for truncation. Rankings reversed across budgets on every benchmark, and the authors argue evaluations should report results conditioned on inference budget. ArXiv · AI/CL/LG's note
score 6