Megadose AI progress, ranked daily.

Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought

· ArXiv · AI/CL/LG ·
The audit found medical chains of thought often did not drive the answer.

Across 14 LLMs and four medical QA benchmarks, destructive clinical edits produced a 72.9% Chain-Decoupling Rate. Corrupting the visible rationale did not change accuracy, and removing CoT prompting did not lower it. Clinician review found 98.5% of 197 perturbed questions still had defensible gold answers. ArXiv · AI/CL/LG's note

score 5

Categories: Research