SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
SoL-Pi cuts token traffic nearly in half while matching Pi-level agent performance on EdgeBench.
The paper frames the gain as coming from harness-level auto-research loops, not a new base model. Four selected mechanisms cover action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench test, it reports a 44.7-49.0% reduction in recorded token traffic and about one-third lower API cost across GPT-5.6 Sol and Opus 5. Estimated savings are listed at $8.75-$13.50 per hour versus native Codex and Claude Code harnesses, and $4.36-$5.71 versus Pi. HF Daily Papers' note
The paper frames the gain as coming from harness-level auto-research loops, not a new base model. Four selected mechanisms cover action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench test, it reports a 44.7-49.0% reduction in recorded token traffic and about one-third lower API cost across GPT-5.6 Sol and Opus 5. Estimated savings are listed at $8.75-$13.50 per hour versus native Codex and Claude Code harnesses, and $4.36-$5.71 versus Pi. HF Daily Papers' note
score 5