Megadose Built for builders and researchers.

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

· HF Daily Papers ·
Recursive Self-Rewrite turns harness-assisted wins into general training data for terminal-task agents.

The paper uses Qwen-3.8-27B to solve tasks under several specialized harnesses, then rewrites successful runs into trajectories that fit a general harness. Its pipeline has a planner, critic, and executor to extract procedures, filter leakage, revise them, and rerun them in fresh sandboxes. Across about 3,000 self-curated terminal tasks, the combined harnesses solved 759 tasks, and RSR expanded 2,001 source trajectories into 11,094 finetuning trajectories. The resulting model beat both the base model and direct trajectory SFT, including a Terminal-Bench 2 pass@3 jump from 57.0% to 74.2%. Source: HF Daily Papers' note.

score 5

Categories: Research