Megadose AI progress, ranked and analyzed.

Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content

· ArXiv · AI/CL/LG ·
Tokens, bytes, and pixels each win under different bottlenecks.

The paper tests the three text encodings on verified parallel sentences in thirteen languages and five scripts, controlling both linguistic content and downstream capacity. Pixels preserve surface form best, bytes are strongest for cross-lingual sentence alignment, and tokens do best on topic classification. The authors argue those results are not reducible to sequence length: short encodings can lose useful meaning, while longer ones can retain information that compresses well. Encoding choice is framed as a task- and capacity-dependent tradeoff, not a universal preference. ArXiv · AI/CL/LG's note

score 4

Categories: Research