Megadose AI progress, ranked and analyzed.

Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot

· ArXiv · AI/CL/LG ·
The paper proposes scoring AI-generated research questions by freezing them in the past and letting later literature judge them.

Hui Mao formalizes “historical backtesting” as a falsifiable alternative to expert scores, LLM judges, and curated examples. The astronomy pilot includes temporally isolated corpora, frozen questions, auditable labels, baselines, and a submission interface. In tests across 2010-2024, evidence-structure-first generation beat LLM-only prompting, while the author says LLM-only systems showed memorized relevance without specific foresight. A released prospective set of 200 questions, frozen on 2026-08-17, is meant to be scored against 2027-2030 literature. ArXiv · AI/CL/LG's note

score 4

Categories: Research