Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
The paper argues that LLM generalization should be measured by how stable a model stays across input variants, not by one aggregate accuracy score.
The authors introduce SAGO, a framework that checks changes in outputs, activations, confidence, and response mirroring across variants of the same input. They say common models show statistically significant instability, with no model generalizing uniformly across the tested axes. The paper also reports that cross-dataset variation can flip model rankings, making single-format benchmark scores misleading. HF Daily Papers' note
The authors introduce SAGO, a framework that checks changes in outputs, activations, confidence, and response mirroring across variants of the same input. They say common models show statistically significant instability, with no model generalizing uniformly across the tested axes. The paper also reports that cross-dataset variation can flip model rankings, making single-format benchmark scores misleading. HF Daily Papers' note
score 4