Megadose AI progress, ranked and analyzed.

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

· ArXiv · AI/CL/LG ·
The benchmark’s best tested model-harness setup reaches only 68% macro-average accuracy, while efficiency varies sharply by harness.

LongHarness Bench is built to separate long-context harnesses that look similar on saturated evaluations. Its tasks make models search through mostly relevant context where only small pieces matter at each reasoning step. The paper tests frontier model families with four state-of-the-art harnesses and finds both accuracy and cost tradeoffs across strategies. The authors frame efficiency as a core evaluation axis, not just a secondary runtime concern. ArXiv · AI/CL/LG's note

score 5

Categories: Research