Self-Supervised Scaling of Terminal Environments for Scientific Domains
The paper turns existing scientific software workflows into verified training tasks for terminal agents.
Its SWR framework runs structured inputs through real workflows, hides some cases for evaluation, and asks an agent to reconstruct an editable program from public examples. A hierarchical verifier checks semantic match, structural validity, and shortcut behavior against hidden outputs. The authors instantiate it with 500 workflows across 46 software families and six domains. Fine-tuning Qwen3.8-27B on the verified reconstructions raises mean Terminal-Bench 2 performance from 47.94% to 53.56%. HF Daily Papers' note
Its SWR framework runs structured inputs through real workflows, hides some cases for evaluation, and asks an agent to reconstruct an editable program from public examples. A hierarchical verifier checks semantic match, structural validity, and shortcut behavior against hidden outputs. The authors instantiate it with 500 workflows across 46 software families and six domains. Fine-tuning Qwen3.8-27B on the verified reconstructions raises mean Terminal-Bench 2 performance from 47.94% to 53.56%. HF Daily Papers' note
score 5