Megadose AI progress, ranked and analyzed.

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

· HF Daily Papers ·
The benchmark tests whether AI systems can learn unfamiliar rules by experimenting, not by recalling training data.

ExplorationBench uses executable “Alien Worlds” where answers can be checked exactly and the rules clash with familiar knowledge. It includes two sandboxes, AlienCode and AlienLogic, with flawed manuals, feedback, and tool-call schemas for agents to explore. The authors report results on 10 AI systems: the strongest can learn and apply new rules, but performance is uneven, and more exploration can sometimes stall or hurt earlier gains. HF Daily Papers' note

score 5

Categories: Research