Megadose AI progress, ranked and analyzed.

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

· HF Daily Papers ·
KoNA tests whether VLMs can answer the parts they should and refuse the parts they should not.

The benchmark covers false premises, visual inaccessibility, universal unknowns, task feasibility, and safety. It evaluates both whole-query refusal and component-level refusal in paired single and compound prompts. The authors report that diverse VLMs fail more often when a query mixes answerable and non-answerable parts. Fine-tuning on KoNA examples improved non-compliance accuracy while largely preserving performance on fully answerable tasks. HF Daily Papers' note

score 5

Categories: Research