R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets
The benchmark finds models often waste shared reasoning budgets even on problems they can solve alone.
R^3-Bench tests six-problem suites across math, programming, and abstract reasoning with one shared budget. Its empirical oracle matched or beat contest performance in all 72 main-table cells, and was strictly better in 71. The authors report weak strategy updating and failures that change with budget pressure. HF Daily Papers' note
R^3-Bench tests six-problem suites across math, programming, and abstract reasoning with one shared budget. Its empirical oracle matched or beat contest performance in all 72 main-table cells, and was strictly better in 71. The authors report weak strategy updating and failures that change with budget pressure. HF Daily Papers' note
score 5