VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
VeriHarness uses the same base model as an agentic verifier, checking both disagreement and consensus before revising long-horizon outputs.
The paper says repeated agent rollouts often surface correct alternatives through disagreement, while shared claims can still be wrong. VeriHarness gives the model a workspace, evidence tools, and reusable verification skills to resolve competing claims and challenge agreed ones. Across five workspace benchmarks and two frontier models, it beat the evaluated selection baselines. Evidence-backed revision improved average performance by 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. HF Daily Papers' note
The paper says repeated agent rollouts often surface correct alternatives through disagreement, while shared claims can still be wrong. VeriHarness gives the model a workspace, evidence tools, and reusable verification skills to resolve competing claims and challenge agreed ones. Across five workspace benchmarks and two frontier models, it beat the evaluated selection baselines. Evidence-backed revision improved average performance by 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. HF Daily Papers' note
score 5