ASI-Bench: At the Dawn of Artificial Superintelligence
AI agents’ research scores fell sharply as ASI-Bench removed human methodological guidance.
The paper introduces ASI-Bench, a 60-task benchmark spanning 11 scientific domains and built with more than 31,000 expert hours. It tests whether systems can choose methods, conduct project-level research, and produce verifiable results as guidance is reduced. Across 18 agent-model setups, average scores dropped from 50.91 with full guidance to 29.10 with only the method specified, and 26.62 when agents had to determine the method themselves. The authors argue this shows current systems remain far from autonomous end-to-end scientific research. HF Daily Papers' note
The paper introduces ASI-Bench, a 60-task benchmark spanning 11 scientific domains and built with more than 31,000 expert hours. It tests whether systems can choose methods, conduct project-level research, and produce verifiable results as guidance is reduced. Across 18 agent-model setups, average scores dropped from 50.91 with full guidance to 29.10 with only the method specified, and 26.62 when agents had to determine the method themselves. The authors argue this shows current systems remain far from autonomous end-to-end scientific research. HF Daily Papers' note
score 5