Megadose AI progress, ranked and analyzed.

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

· ArXiv · AI/CL/LG ·
The paper says LLM judges form summary ratings through a two-stage internal pipeline, not just a final score heuristic.

The authors test Themis and Prometheus with controlled corruptions to readability and adequacy. They find lower layers route local error evidence, while later MLP layers integrate that signal and write the rating. The rating appears to settle sharply near layer 26 for Themis and layer 25 for Prometheus. A base Llama-3-8B model shows part of the same routing structure, suggesting fine-tuning reshapes an existing mechanism rather than creating it from scratch. ArXiv · AI/CL/LG's note

score 5

Categories: Research