D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models
The paper proposes a one-pass hidden-state score that flags hallucinations without retrieval, external verification, or repeated generations.
D-Score counts how many singular directions in a model’s hidden activations stay close to the leading singular value. The authors treat a higher count as a hallucination signal, arguing that conflicting or unsupported content can spread the hidden trajectory across more directions. They evaluate it on FAVA-Annotation and RAGTruth and report that it performs as a strong hidden-state detector. ArXiv · AI/CL/LG's note
D-Score counts how many singular directions in a model’s hidden activations stay close to the leading singular value. The authors treat a higher count as a hallucination signal, arguing that conflicting or unsupported content can spread the hidden trajectory across more directions. They evaluate it on FAVA-Annotation and RAGTruth and report that it performs as a strong hidden-state detector. ArXiv · AI/CL/LG's note
score 4