TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
TasteVal claims the strongest tested model beat the best human expert baseline on experimental compute efficiency.
The benchmark tests how well models choose experiments and interpret results on fixed AI R&D problems, while a separate coder agent handles implementation. Its authors report Opus 5.5 at a 2.3x compute multiplier over the expert baseline, with a 95% confidence interval of 1.15-4.37. The paper says frontier-model compute multipliers on TasteVal have doubled about every 3.0 months since December 2025. The tasks are not being released to avoid contamination. ArXiv · AI/CL/LG's note
The benchmark tests how well models choose experiments and interpret results on fixed AI R&D problems, while a separate coder agent handles implementation. Its authors report Opus 5.5 at a 2.3x compute multiplier over the expert baseline, with a 95% confidence interval of 1.15-4.37. The paper says frontier-model compute multipliers on TasteVal have doubled about every 3.0 months since December 2025. The tasks are not being released to avoid contamination. ArXiv · AI/CL/LG's note
score 6