LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning
The benchmark’s best tested model-harness setup reaches only 68% macro-average accuracy, while efficiency varies sharply by harness.
LongHarness Bench is built to separate long-context harnesses that look similar on saturated evaluations. Its tasks make models search through mostly relevant context where only small pieces matter at each reasoning step. The paper tests frontier model families with four state-of-the-art harnesses and finds both accuracy and cost tradeoffs across strategies. The authors frame efficiency as a core evaluation axis, not just a secondary runtime concern. ArXiv · AI/CL/LG's note
LongHarness Bench is built to separate long-context harnesses that look similar on saturated evaluations. Its tasks make models search through mostly relevant context where only small pieces matter at each reasoning step. The paper tests frontier model families with four state-of-the-art harnesses and finds both accuracy and cost tradeoffs across strategies. The authors frame efficiency as a core evaluation axis, not just a secondary runtime concern. ArXiv · AI/CL/LG's note
score 5