ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Author-labeled “inspiration” retrieval is still hard for current AI search agents.
ScholarCatalyst is built from 184 lead authors of 207 recent CS papers, who marked earlier papers that did or could have advanced their projects and gave rationales. The task asks systems to retrieve those papers from the literature available when each project began. Agentic search trails embedding retrieval at Recall@20, and a Claude Fable 5.1-based agent reaches only 0.51. The authors frame the benchmark as a way to train models toward expert intuition for finding prior work behind new research ideas. ArXiv · AI/CL/LG's note
ScholarCatalyst is built from 184 lead authors of 207 recent CS papers, who marked earlier papers that did or could have advanced their projects and gave rationales. The task asks systems to retrieve those papers from the literature available when each project began. Agentic search trails embedding retrieval at Recall@20, and a Claude Fable 5.1-based agent reaches only 0.51. The authors frame the benchmark as a way to train models toward expert intuition for finding prior work behind new research ideas. ArXiv · AI/CL/LG's note
score 4