When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence
The paper says robots should decide whether to ask for help from measured diagnostic accuracy and cost, not model confidence.
The authors built a simulated benchmark where robot failures had known injected causes and audited which sensors could reveal them. Some causes appeared in images, while others were only diagnosable from force data. Six open vision-language models often followed prompt wording instead of evidence, including sharp shifts in refusal rates when answer order changed. Giving force data as text produced above-baseline diagnoses in four of six models, but their ask rates still failed to track the decision policy. ArXiv · AI/CL/LG's note
The authors built a simulated benchmark where robot failures had known injected causes and audited which sensors could reveal them. Some causes appeared in images, while others were only diagnosable from force data. Six open vision-language models often followed prompt wording instead of evidence, including sharp shifts in refusal rates when answer order changed. Giving force data as text produced above-baseline diagnoses in four of six models, but their ask rates still failed to track the decision policy. ArXiv · AI/CL/LG's note
score 5