Megadose AI progress, ranked and analyzed.

Aftermarket Harnesses

Tomasz Tunguz ·
The same model scored sharply differently depending on the coding harness around it.

Tunguz cites Endor Labs results showing GPT-5.5 at 61.5% functional correctness in Codex and 87.2% in Cursor. Claude Opus 4.7 also scored higher in Cursor than in Claude Code. The note argues that harness choices around context, retrieval, and prompt caching now shape both benchmark performance and cost. Input tokens dominate OpenRouter volume, making cache discipline a central part of the bill. Tomasz Tunguz's note

score 5

Categories: Research