DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat
Frontier models can handle single D&D fights, but stumble when combat choices carry across an adventuring day.
DungeonBench tests tactical reasoning with legal, simulator-generated options covering movement, attacks, spells, reactions, objectives, preparation, and scarce resources. It has an Encounter track for single fights and a Day track that preserves hit points, spell slots, consumables, preparation, and rest timing across linked encounters. The authors report that full tactical observations do not saturate the benchmark: strong language-model policies often win individual encounters, while the linked track exposes weaker resource budgeting and rule-aware discipline. ArXiv · AI/CL/LG's note
DungeonBench tests tactical reasoning with legal, simulator-generated options covering movement, attacks, spells, reactions, objectives, preparation, and scarce resources. It has an Encounter track for single fights and a Day track that preserves hit points, spell slots, consumables, preparation, and rest timing across linked encounters. The authors report that full tactical observations do not saturate the benchmark: strong language-model policies often win individual encounters, while the linked track exposes weaker resource budgeting and rule-aware discipline. ArXiv · AI/CL/LG's note
score 5