Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
The paper makes “skill switching” the thing being measured, then uses it as a training reward.
The authors define Skill Entropy as a way to score how hard it is for a model to move between reasoning skills inside one long task. They introduce Skill²-Bench, covering 558 skills across 9 domains, and report that accuracy drops as task entropy rises. Their Skill-Entropy RL method trains models to predict both answers and the skill sequence behind them. On two Qwen3 models, the reported benchmark scores rise from 34.4% to 68.4% and from 14.6% to 40.1%. ArXiv · AI/CL/LG's note
The authors define Skill Entropy as a way to score how hard it is for a model to move between reasoning skills inside one long task. They introduce Skill²-Bench, covering 558 skills across 9 domains, and report that accuracy drops as task entropy rises. Their Skill-Entropy RL method trains models to predict both answers and the skill sequence behind them. On two Qwen3 models, the reported benchmark scores rise from 34.4% to 68.4% and from 14.6% to 40.1%. ArXiv · AI/CL/LG's note
score 5