A Formal Limitation on Learning Human Language From Textual Corpora
Utterances alone leave some intended meaning unrecoverable, no matter how strong the text representation is.
Cheng and Cotterell frame language use as a distribution over meanings, contexts, and utterances, then bound how well any decoder can recover a speaker’s meaning from the utterance representation. The limit applies even to features drawn from contemporary LLM hidden states. Their account separates uncertainty that is irreducible from uncertainty that extralinguistic context can resolve but text alone cannot. They report supporting experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference. ArXiv · AI/CL/LG's note
Cheng and Cotterell frame language use as a distribution over meanings, contexts, and utterances, then bound how well any decoder can recover a speaker’s meaning from the utterance representation. The limit applies even to features drawn from contemporary LLM hidden states. Their account separates uncertainty that is irreducible from uncertainty that extralinguistic context can resolve but text alone cannot. They report supporting experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference. ArXiv · AI/CL/LG's note
score 5