ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
ChronoLens compares language change in one shared space across five parliamentary corpora from 1803 to 2026.
The paper applies frozen multilingual language models, feature-aligned crosscoders, and linguistic interventions to 44.98 million documents, about 17.2 billion tokens. Its sparse representations track linguistic statistics more closely than dense embeddings or a pooled sparse autoencoder. The authors find that morphology, syntax, semantics, and pragmatics tend to shift by similar amounts within a language, while different languages change at different times, distances, and directions. HF Daily Papers' note
The paper applies frozen multilingual language models, feature-aligned crosscoders, and linguistic interventions to 44.98 million documents, about 17.2 billion tokens. Its sparse representations track linguistic statistics more closely than dense embeddings or a pooled sparse autoencoder. The authors find that morphology, syntax, semantics, and pragmatics tend to shift by similar amounts within a language, while different languages change at different times, distances, and directions. HF Daily Papers' note
score 4