Megadose Built for builders and researchers.

TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts

· ArXiv · AI/CL/LG ·
TasteVal claims the strongest tested model beat the best human expert baseline on experimental compute efficiency.

The benchmark tests how well models choose experiments and interpret results on fixed AI R&D problems, while a separate coder agent handles implementation. Its authors report Opus 5.5 at a 2.3x compute multiplier over the expert baseline, with a 95% confidence interval of 1.15-4.37. The paper says frontier-model compute multipliers on TasteVal have doubled about every 3.0 months since December 2025. The tasks are not being released to avoid contamination. ArXiv · AI/CL/LG's note

score 6

Categories: Research