Megadose Built for builders and researchers.

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

· ArXiv · AI/CL/LG ·
StrategyBench tests whether models can turn few-shot examples into explicit task rules that still help on new inputs.

The paper builds the benchmark from strategy-inducible BIG-Bench tasks, with reference strategies and metrics for both strategy quality and downstream utility. The authors compare how results change by task category, generator-executor setup, demonstration design, and SFT-based adaptation. Their experiments find that explicit strategies help unevenly: usefulness depends heavily on both the task and the conditions under which the strategy is generated and executed. ArXiv · AI/CL/LG's note

score 4

Categories: Research