Megadose AI progress, ranked and analyzed.

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

· ArXiv · AI/CL/LG ·
The benchmark turns working web apps into replay-checkable coding tasks.

ProgramDistill has agents infer missing behavior by interacting with complete reference applications, then implement it in incomplete versions. The authors report 1,975 replay-verified behaviors mined from 26 apps, producing 4,063 tasks without human intervention. In their tests across nine frontier coding agents, GPT-6 Astra reached 49.2% success on cumulative full-app reconstruction, while Claude Opus 5 reached 28.8%. Performance dropped sharply as partial reconstruction required deeper restoration. ArXiv · AI/CL/LG's note

score 6

Categories: Research