Megadose AI progress, ranked and analyzed.

Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

· ArXiv · AI/CL/LG ·
The paper argues that evaluator behavior needs controlled stress tests, not just aggregate agreement scores.

The authors propose “behavioral correctness assumptions” for reference-based automatic evaluation methods in natural language generation. They test lexical, character-level, semantic, LLM-based, and hybrid evaluators with controlled response transformations and expected scoring behavior. The reported result is that no evaluator satisfies all of the proposed assumptions. Evaluators with similar aggregate performance can behave very differently under these assumption-level checks. ArXiv · AI/CL/LG's note

score 4

Categories: Research