Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity
The paper finds LLMs can sense when they do not know a referent, but still generate overly specific answers.
The authors test models on a T-REx-based benchmark varying entity familiarity and specificity. They report that model activations encode both knowledge-boundary status and expected referent specificity. Those signals do not appear to control generation: models still favor specific referents, even for unknown entities and even when correct generic alternatives are available. The paper frames this as groundwork for “Gricean alignment,” linking uncertainty to safer, less specific output. ArXiv · AI/CL/LG's note
The authors test models on a T-REx-based benchmark varying entity familiarity and specificity. They report that model activations encode both knowledge-boundary status and expected referent specificity. Those signals do not appear to control generation: models still favor specific referents, even for unknown entities and even when correct generic alternatives are available. The paper frames this as groundwork for “Gricean alignment,” linking uncertainty to safer, less specific output. ArXiv · AI/CL/LG's note
score 5