TestPrism: Rethinking Test Evaluation Beyond a Single Reference
TestPrism says single-reference grading may nearly double measured test success.
The paper introduces a benchmark with 300 tasks and 3,000 candidate implementations, split between valid and invalid solutions. Its Joint Success Function requires generated tests to catch the initial bug, accept all valid alternatives, and reject all invalid ones. Fourteen baseline coding-agent setups reached 28.00% on that metric, versus 59.67% under single-reference success. The authors also propose TestHelix, which improved Joint Success Function by 8.67 to 9.00 percentage points across two models. HF Daily Papers' note
The paper introduces a benchmark with 300 tasks and 3,000 candidate implementations, split between valid and invalid solutions. Its Joint Success Function requires generated tests to catch the initial bug, accept all valid alternatives, and reject all invalid ones. Fourteen baseline coding-agent setups reached 28.00% on that metric, versus 59.67% under single-reference success. The authors also propose TestHelix, which improved Joint Success Function by 8.67 to 9.00 percentage points across two models. HF Daily Papers' note
score 4