Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
The audit found medical chains of thought often did not drive the answer.
Across 14 LLMs and four medical QA benchmarks, destructive clinical edits produced a 72.9% Chain-Decoupling Rate. Corrupting the visible rationale did not change accuracy, and removing CoT prompting did not lower it. Clinician review found 98.5% of 197 perturbed questions still had defensible gold answers. ArXiv · AI/CL/LG's note
Across 14 LLMs and four medical QA benchmarks, destructive clinical edits produced a 72.9% Chain-Decoupling Rate. Corrupting the visible rationale did not change accuracy, and removing CoT prompting did not lower it. Clinician review found 98.5% of 197 perturbed questions still had defensible gold answers. ArXiv · AI/CL/LG's note
score 5