Does task decomposition improve automatic NLG evaluation?
The paper finds no performance gain from decomposing LLM-judge evaluations once the baseline is controlled.
The authors compare decomposed and non-decomposed LLMaJ methods across multiple NLG datasets. They argue earlier gains came from using human labels as training data, not from decomposition itself. When those labels are available, a non-decomposed LLMaJ setup can perform comparably to human annotators. ArXiv · AI/CL/LG's note
The authors compare decomposed and non-decomposed LLMaJ methods across multiple NLG datasets. They argue earlier gains came from using human labels as training data, not from decomposition itself. When those labels are available, a non-decomposed LLMaJ setup can perform comparably to human annotators. ArXiv · AI/CL/LG's note
score 5