Megadose Built for builders and researchers.

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

· HF Daily Papers ·
The benchmark tests whether AI can find the earlier papers that would have helped a research project before it was finished.

ScholarCatalyst is built from 184 lead authors of 207 recent computer science papers, who labeled candidate prior papers and gave rationales. The task asks systems to retrieve those useful papers using only the literature available when the project began. Agentic search trails or roughly matches embedding retrieval, with reported Recall@20 scores of 0.42 versus 0.48. A Claude Fable 5.1-based agent reaches 0.51, still leaving the authors arguing for training methods that better capture expert research intuition. HF Daily Papers' note

score 5

Categories: Research