ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step
The benchmark finds agents can discover hidden tools, then fail to adapt when those tools change.
ScrambleToolBench strips away semantic tool cues and forces agents to infer behavior through terminal interaction. The paper adds mapping drift, stochastic failures, and timed execution windows to test whether agents revise their hypotheses. In the evaluation, state-of-the-art models often fall back to exhaustive search instead of deductive recovery, even when more test-time reasoning is allowed. Persistent memory reduces compounding errors, but does not solve efficient adaptation. HF Daily Papers' note
ScrambleToolBench strips away semantic tool cues and forces agents to infer behavior through terminal interaction. The paper adds mapping drift, stochastic failures, and timed execution windows to test whether agents revise their hypotheses. In the evaluation, state-of-the-art models often fall back to exhaustive search instead of deductive recovery, even when more test-time reasoning is allowed. Persistent memory reduces compounding errors, but does not solve efficient adaptation. HF Daily Papers' note
score 5