Megadose AI progress, ranked and analyzed.

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

· HF Daily Papers ·
Taste-Bench tests whether an agent can pick the better path before seeing how the run ends.

The paper builds decision-fork questions from engineering and research agent trajectories, without human annotation. Frontier models top out at 59.7% accuracy, and later-arriving evidence makes the choices harder. More reasoning budget does not improve results. The authors also report that distilling outcome-aware judgments into a student model improves both fork decisions and held-out SWE-bench Pro success. HF Daily Papers' note

score 5

Categories: Research