Megadose AI progress, ranked and analyzed.

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

· HF Daily Papers ·
GAMUT turns structured factual rubrics into binary checks so LLM judges can score long-form answers more consistently.

The paper frames this as a way around the tradeoff between expressive rubrics and reliable grading. Its benchmark contains 1,813 questions grounded in real wearable imagery across 10 domains, with expert-verified, evidence-backed rubrics. In tests on 14 frontier and open-weight models, the best reported score was 58.7% from Gemini 3.1 Pro, suggesting the benchmark remains difficult and separates model performance. HF Daily Papers' note

score 5

Categories: Research