Evaluation of Contextual Understanding in Large Language Models
The paper proposes a knowledge-graph metric for testing whether LLM answers are grounded in context.
The authors argue that perplexity, BLEU, and surface accuracy miss whether a model actually extracts and reasons over supplied context. Their framework introduces Semantic Structural Similarity for KGs, or S3KG, combining semantic and structural comparison into a continuous score. They also add a diagnostic layer for categorizing reasoning errors in model answers. The validation is on a curated QA benchmark, where they compare S3KG with established metrics for correctness, faithfulness, and interpretability. ArXiv · AI/CL/LG's note
The authors argue that perplexity, BLEU, and surface accuracy miss whether a model actually extracts and reasons over supplied context. Their framework introduces Semantic Structural Similarity for KGs, or S3KG, combining semantic and structural comparison into a continuous score. They also add a diagnostic layer for categorizing reasoning errors in model answers. The validation is on a curated QA benchmark, where they compare S3KG with established metrics for correctness, faithfulness, and interpretability. ArXiv · AI/CL/LG's note
score 4