Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
Argus is pitched as a fixed-weight agent runtime that improves through persistent state, verification, and reviewed control decisions.
The paper describes a Manager-Planner-Engineer-Reviewer setup for long-horizon tasks, with durable project memory and escalation points for operators. On seven GPT-5.5 benchmark arenas, it reports about 78% on SWE-Bench Pro, versus 59% for Direct Copilot, at 1.41x aggregate tokens. After verification-gated self-evolution, later SWE-Bench waves used fewer solve-input tokens and less active workflow time than startup waves. The authors also cite an upstream RWKV6 kernel merge, a multi-day math campaign, and six paper pipelines as non-benchmark runs. ArXiv · AI/CL/LG's note
The paper describes a Manager-Planner-Engineer-Reviewer setup for long-horizon tasks, with durable project memory and escalation points for operators. On seven GPT-5.5 benchmark arenas, it reports about 78% on SWE-Bench Pro, versus 59% for Direct Copilot, at 1.41x aggregate tokens. After verification-gated self-evolution, later SWE-Bench waves used fewer solve-input tokens and less active workflow time than startup waves. The authors also cite an upstream RWKV6 kernel merge, a multi-day math campaign, and six paper pipelines as non-benchmark runs. ArXiv · AI/CL/LG's note
score 5