Megadose AI progress, ranked and analyzed.

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications

· ArXiv · AI/CL/LG ·
The benchmark finds LLMs can ask useful clarifying questions once a defect is identified, but mostly fail to locate the missing implementation detail themselves.

IdeaAMBIG contains 660 evidence-grounded cases drawn from reproducibility reports, GitHub issues, papers, codebases, and artifacts. Across 13 models, the best result on real-world defect recovery was 9.6% Macro Defect Recovery Rate. When the annotated defect was supplied, the best clarification-action score rose to 80.6%. An oracle resolution study raised codification readiness from 14% to 98%, pointing to localization as the main bottleneck. ArXiv · AI/CL/LG's note

score 4

Categories: Research