Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
GAMUT turns structured factual rubrics into binary checks so LLM judges can score long-form answers more consistently.
The paper frames this as a way around the tradeoff between expressive rubrics and reliable grading. Its benchmark contains 1,813 questions grounded in real wearable imagery across 10 domains, with expert-verified, evidence-backed rubrics. In tests on 14 frontier and open-weight models, the best reported score was 58.7% from Gemini 3.1 Pro, suggesting the benchmark remains difficult and separates model performance. HF Daily Papers' note
The paper frames this as a way around the tradeoff between expressive rubrics and reliable grading. Its benchmark contains 1,813 questions grounded in real wearable imagery across 10 domains, with expert-verified, evidence-backed rubrics. In tests on 14 frontier and open-weight models, the best reported score was 58.7% from Gemini 3.1 Pro, suggesting the benchmark remains difficult and separates model performance. HF Daily Papers' note
score 5