Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models
The paper argues that names are not equal inputs for LLMs before any prompt behavior is measured.
Some names are represented as single tokens, while others are split into subwords, and that support varies across tokenizers. The authors test nearly half a million first names across 12 LLM-associated tokenizers and find uneven access tied to race- and gender-associated metadata. Their NameTrace framework measures whether that tokenizer-level difference shows up inside task-relevant representations. In matched comparisons, lexical support predicts differences across fellowship, hiring, clinical assessment, and lending settings, with hidden-state interventions showing those directions can affect later constrained choices. HF Daily Papers' note
Some names are represented as single tokens, while others are split into subwords, and that support varies across tokenizers. The authors test nearly half a million first names across 12 LLM-associated tokenizers and find uneven access tied to race- and gender-associated metadata. Their NameTrace framework measures whether that tokenizer-level difference shows up inside task-relevant representations. In matched comparisons, lexical support predicts differences across fellowship, hiring, clinical assessment, and lending settings, with hidden-state interventions showing those directions can affect later constrained choices. HF Daily Papers' note
score 4