The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
Taste-Bench tests whether an agent can pick the better path before seeing how the run ends.
The paper builds decision-fork questions from engineering and research agent trajectories, without human annotation. Frontier models top out at 59.7% accuracy, and later-arriving evidence makes the choices harder. More reasoning budget does not improve results. The authors also report that distilling outcome-aware judgments into a student model improves both fork decisions and held-out SWE-bench Pro success. HF Daily Papers' note
The paper builds decision-fork questions from engineering and research agent trajectories, without human annotation. Frontier models top out at 59.7% accuracy, and later-arriving evidence makes the choices harder. More reasoning budget does not improve results. The authors also report that distilling outcome-aware judgments into a student model improves both fork decisions and held-out SWE-bench Pro success. HF Daily Papers' note
score 5