Megadose AI progress, ranked and analyzed.

LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora

· ArXiv · AI/CL/LG ·
The audit finds LLM essay graders can be consistent with themselves while still shifting sharply by model, provider, and version.

Across 2,377 public-corpus essays, judge severity varied by 219 points on ENEM’s 0–1000 scale and by 15–33% of the ASAP score range. Judge-human correlations clustered at .47–.56, which the paper describes as undiscriminating. All five version comparisons showed severity shifts beyond the permutation null, with changes up to 133 points. The study reports no credible evidence that LLM halo exceeded the trained-human range after a same-instrument check. ArXiv · AI/CL/LG's note

score 4

Categories: Research