Megadose AI progress, ranked and analyzed.

Local and Global Regimes of Geometric Complexity in Language Model Representations

· ArXiv · AI/CL/LG ·
Lexical diversity can flip intrinsic-dimensionality readings, making a common complexity probe dataset-sensitive.

The paper tests how the number of unique last-token words in a dataset affects intrinsic dimensionality estimates in language model representations. It reports two scale-dependent regimes: at low lexical diversity, fewer unique final words yield higher ID; at high lexical diversity, more unique words do. The authors derive an exact, parameter-free formula for the reversal point and say it matches every tested scale. Their caution is narrow but sharp: ID should not be read as a straightforward measure of representational complexity without accounting for dataset construction. ArXiv · AI/CL/LG's note

score 4

Categories: Research