Mapping and Measuring the Behavioral Evolution of Large Language Models
The paper compares 32 language models by their outputs on the same 10,000 prompts, treating behavior as something to map across model families and generations.
The authors build three sentence-level distance measures and use them to track clustering, drift, convergence, and response dispersion. They report that model families form coherent behavioral clusters, with `gpt-2` standing apart as a global outlier. Cross-family distances decline over time, and several recent reasoning-oriented models show more compact response clouds. A token-level check closely matches the sentence-level result, with Spearman `ρ=0.98`. ArXiv · AI/CL/LG's note
The authors build three sentence-level distance measures and use them to track clustering, drift, convergence, and response dispersion. They report that model families form coherent behavioral clusters, with `gpt-2` standing apart as a global outlier. Cross-family distances decline over time, and several recent reasoning-oriented models show more compact response clouds. A token-level check closely matches the sentence-level result, with Spearman `ρ=0.98`. ArXiv · AI/CL/LG's note
score 5