WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution
Candidate senses made WiC easier for open LLMs across every setting tested.
The paper argues that WiC’s difficulty is partly about missing sense inventories, not just comparing two uses of a word. When models were given candidate senses in a WSD-like setup, their judgments became more consistent and targeted. The authors also found that some apparent WiC mistakes came from ambiguous labels or mismatched sense boundaries between models and annotators. One recurring failure mode was that LLMs drew overly fine-grained distinctions.
ArXiv · AI/CL/LG's note
The paper argues that WiC’s difficulty is partly about missing sense inventories, not just comparing two uses of a word. When models were given candidate senses in a WSD-like setup, their judgments became more consistent and targeted. The authors also found that some apparent WiC mistakes came from ambiguous labels or mismatched sense boundaries between models and annotators. One recurring failure mode was that LLMs drew overly fine-grained distinctions.
ArXiv · AI/CL/LG's note
score 4