Aftermarket Harnesses
Benchmark results show coding-agent harnesses can dominate model performance and token costs through context management and caching.
Excerpt
The harness now moves the coding benchmark more than the model does. Endor Labs' Agent Security League found GPT-5.5 scored 61.5% functional correctness in Codex & 87.2% in Cursor, & Claude Opus 4.7 scored 87.2% in Claude Code & 91.1% in Cursor. Input tokens are 86-98% of OpenRouter volume, so the harness controls most of the bill through cache discipline. First-party co-design buys real cache hit rates, but a third-party harness can match them.
Read at source: https://www.tomtunguz.com/aftermarket-harnesses/