Megadose Built for builders and researchers.

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

· HF Daily Papers ·
Verified prior execution helped models finish more long workflows, not just summarize them better.

The paper introduces LongWoF-Bench, a 778-task benchmark with machine-verifiable outcomes across coding, agent-environment synthesis, math reasoning, and rule following. On 252 tasks with verifier-confirmed Claude Opus trajectories, EvoMap “Gene” artifacts beat Skill baselines across seven models by 8.7 to 15.5 points. The authors say reference-distilled Gene did not show the same edge, tying the gains to verified execution provenance. For Claude Opus, Gene reuse solved 39 more tasks than Skill while cutting solve-time token use by 9.9%. HF Daily Papers' note

score 5

Categories: Research