Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station
Station agents recovered 62.7% of withheld findings across three open-ended research tasks.
The paper tests whether agents can make progress when they are given only a research question, not target metrics or the original results. Station adds a Supervisor mechanism and periodic Meta Reflection to keep exploration going when intermediate feedback is thin. In the authors’ benchmark, it outperformed Codex Multiagent-v2 and AI Scientist-v2 by a wide margin. The paper also reports two non-oracle tasks where agent discoveries closely matched later researcher findings. Source: HF Daily Papers' note.
The paper tests whether agents can make progress when they are given only a research question, not target metrics or the original results. Station adds a Supervisor mechanism and periodic Meta Reflection to keep exploration going when intermediate feedback is thin. In the authors’ benchmark, it outperformed Codex Multiagent-v2 and AI Scientist-v2 by a wide margin. The paper also reports two non-oracle tasks where agent discoveries closely matched later researcher findings. Source: HF Daily Papers' note.
score 5