TokEval: A Tokenizer Evaluation Suite
Tokenizer choices are tested here as measurable causes of model behavior, not setup trivia.
Clara Meister introduces TokEval, a suite of intrinsic tokenizer metrics covering properties such as UTF-8 boundary integrity, digit place-value alignment, fertility, and compression. The paper validates those metrics with controlled pretraining runs where only tokenizer data mix, pretokenization, and training algorithm vary. Information-theoretic metrics correlate with bits-per-byte language modeling results, while structure-sensitive measures line up with accuracy on tasks like math and code. ArXiv · AI/CL/LG's note
Clara Meister introduces TokEval, a suite of intrinsic tokenizer metrics covering properties such as UTF-8 boundary integrity, digit place-value alignment, fertility, and compression. The paper validates those metrics with controlled pretraining runs where only tokenizer data mix, pretokenization, and training algorithm vary. Information-theoretic metrics correlate with bits-per-byte language modeling results, while structure-sensitive measures line up with accuracy on tasks like math and code. ArXiv · AI/CL/LG's note
score 5