Megadose AI progress, ranked and analyzed.

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications

· HF Daily Papers ·
The benchmark finds LLMs can often ask the right clarification once a defect is identified, but mostly fail to locate the defect themselves.

IdeaAMBIG tests whether research-method specs are ready to be implemented without unsupported assumptions. It includes 660 evidence-grounded cases, drawn from reproducibility reports, GitHub issues, papers, codebases, and controlled synthetic gaps. Across 13 models, the best result on real-world defect recovery was 9.6%, while clarification-action success reached 80.6% when the defect was already given. An oracle resolution study raised downstream codification readiness from 14% to 98%. HF Daily Papers' note

score 5

Categories: Research