Megadose Built for builders and researchers.

Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification

· HF Daily Papers ·
Bigger models and stronger retrieval did not fix the failure patterns in open-ended medical fact checks.

The paper studies retrieval-based factuality evaluation for long-form medical answers, where generated claims are checked against authoritative medical evidence. It builds taxonomies for retrieval errors and verifier reasoning errors, then labels those failures at scale with an LLM-as-judge pipeline. The authors test across multiple retrieval methods and frontier verifier models, and report that model scaling, extra reasoning effort, web sources, and medical fine-tuning still leave the same failure modes. HF Daily Papers' note

score 4

Categories: Research