Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
KoNA tests whether VLMs can answer the parts they should and refuse the parts they should not.
The benchmark covers false premises, visual inaccessibility, universal unknowns, task feasibility, and safety. It evaluates both whole-query refusal and component-level refusal in paired single and compound prompts. The authors report that diverse VLMs fail more often when a query mixes answerable and non-answerable parts. Fine-tuning on KoNA examples improved non-compliance accuracy while largely preserving performance on fully answerable tasks. HF Daily Papers' note
The benchmark covers false premises, visual inaccessibility, universal unknowns, task feasibility, and safety. It evaluates both whole-query refusal and component-level refusal in paired single and compound prompts. The authors report that diverse VLMs fail more often when a query mixes answerable and non-answerable parts. Fine-tuning on KoNA examples improved non-compliance accuracy while largely preserving performance on fully answerable tasks. HF Daily Papers' note
score 5