Temporal Multi-Signal Fusion for Token-Level Hallucination Detection
The paper’s bet is that hallucinations are spans, not isolated token mistakes.
Igor Itkin’s detector scores each token using a 33-feature stream combining text statistics, NLI entailment, and language-model surprisal, without needing access to model internals. A BiGRU sequence labeler reaches 0.840 AUC on RAGTruth, beating an independent logistic-regression baseline by 11 points. The paper attributes most of that gain to temporal ordering, where evidence from clearer tokens helps classify nearby ambiguous ones. Similar ceilings across BiGRU, Mamba, and attention models suggest the limiting factor is the feature set. HF Daily Papers' note
Igor Itkin’s detector scores each token using a 33-feature stream combining text statistics, NLI entailment, and language-model surprisal, without needing access to model internals. A BiGRU sequence labeler reaches 0.840 AUC on RAGTruth, beating an independent logistic-regression baseline by 11 points. The paper attributes most of that gain to temporal ordering, where evidence from clearer tokens helps classify nearby ambiguous ones. Similar ceilings across BiGRU, Mamba, and attention models suggest the limiting factor is the feature set. HF Daily Papers' note
score 5