The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
The paper argues that hallucination benchmarks can change meaning depending on how their answers are labeled.
The authors test 900 human-labeled QA pairs across three datasets and three generator models. They find large disagreements among lexical metrics, NLI, and seven LLM judges, as well as between those automated labels and human factual-correctness annotations. Prompts aimed at factual correctness usually align better with humans than faithfulness-style prompts. Their conclusion is that label source and labeling criterion need to be explicit parts of benchmark design. ArXiv · AI/CL/LG's note
The authors test 900 human-labeled QA pairs across three datasets and three generator models. They find large disagreements among lexical metrics, NLI, and seven LLM judges, as well as between those automated labels and human factual-correctness annotations. Prompts aimed at factual correctness usually align better with humans than faithfulness-style prompts. Their conclusion is that label source and labeling criterion need to be explicit parts of benchmark design. ArXiv · AI/CL/LG's note
score 5