LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
The paper’s central move is to take task state out of the agent’s running context and update it only after external verification.
LongHorizon-Harness uses a Manage-Execute-Audit loop: a manager tracks state and picks the next subtask, a fresh-context executor acts, and a read-only auditor checks the environment before progress is recorded. The authors argue this limits cascading errors from confused state tracking or premature self-assessment. Reported gains include Qwen 3.7-Plus rising from 51.8% to 80.7% on WeaveBench and Claude Opus 4.7 improving from 20.0% to 34.3% on an OSWorld 2.0 subset. Source: HF Daily Papers' note.
LongHorizon-Harness uses a Manage-Execute-Audit loop: a manager tracks state and picks the next subtask, a fresh-context executor acts, and a read-only auditor checks the environment before progress is recorded. The authors argue this limits cascading errors from confused state tracking or premature self-assessment. Reported gains include Qwen 3.7-Plus rising from 51.8% to 80.7% on WeaveBench and Claude Opus 4.7 improving from 20.0% to 34.3% on an OSWorld 2.0 subset. Source: HF Daily Papers' note.
score 5