Megadose AI progress, ranked and analyzed.

Can AI agents conduct open-ended AI research? Early evidence from two case studies

· HF Daily Papers ·
Frontier agents finished the coding work, but the paper’s authors rejected both research outputs as not real progress.

The authors tested “shadow evaluations” on two unpublished NeurIPS 2026 submissions, giving agents six days and substantial compute. In both cases, the agents handled the engineering without human help but failed to answer the open-ended research questions in a publishable way. The paper names recurring failures in judgment, research design, backtracking, resource awareness, and instruction drift. A second model and scaffold showed the same pattern. HF Daily Papers' note

score 5

Categories: Research