ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
ProgramDistill tests coding agents on rebuilding app behavior learned from working reference software.
The benchmark turns interactive web apps into verifiable software-engineering tasks by mining replayable behaviors from gold patches. Its pipeline found 1,975 verified behaviors across 26 applications and generated 4,063 tasks without human intervention. In full-application reconstruction, GPT-6 Astra scored 49.2% on cumulative workflows, while Claude Opus 5 scored 28.8%. Performance dropped as partial reconstruction got deeper, giving the benchmark a controlled difficulty scale. HF Daily Papers' note
The benchmark turns interactive web apps into verifiable software-engineering tasks by mining replayable behaviors from gold patches. Its pipeline found 1,975 verified behaviors across 26 applications and generated 4,063 tasks without human intervention. In full-application reconstruction, GPT-6 Astra scored 49.2% on cumulative workflows, while Claude Opus 5 scored 28.8%. Performance dropped as partial reconstruction got deeper, giving the benchmark a controlled difficulty scale. HF Daily Papers' note
score 6