DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Longer chat histories made models more likely to show delusion-linked behavior.
The paper introduces DelusionEval, built from 589 conversation histories involving users who experienced delusions and psychological harm. Across model families, the authors found substantial rates of behavior they associate with promoting delusions. Model size, release date, and test-time reasoning did not reliably predict safer behavior. Adding 350 prior messages raised one self-harm-related failure rate from 30.0% to 41.1%. ArXiv · AI/CL/LG's note
The paper introduces DelusionEval, built from 589 conversation histories involving users who experienced delusions and psychological harm. Across model families, the authors found substantial rates of behavior they associate with promoting delusions. Model size, release date, and test-time reasoning did not reliably predict safer behavior. Adding 350 prior messages raised one self-harm-related failure rate from 30.0% to 41.1%. ArXiv · AI/CL/LG's note
score 5