Megadose AI progress, ranked and analyzed.

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

· HF Daily Papers ·
ProgramDistill tests coding agents on rebuilding app behavior learned from working reference software.

The benchmark turns interactive web apps into verifiable software-engineering tasks by mining replayable behaviors from gold patches. Its pipeline found 1,975 verified behaviors across 26 applications and generated 4,063 tasks without human intervention. In full-application reconstruction, GPT-6 Astra scored 49.2% on cumulative workflows, while Claude Opus 5 scored 28.8%. Performance dropped as partial reconstruction got deeper, giving the benchmark a controlled difficulty scale. HF Daily Papers' note

score 6

Categories: Research