Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
The paper says existing scaling laws miss at the data-scarce and overtraining edges because they treat model size and data as independent.
The authors propose “Skaling,” which adds a single interaction exponent coupling capacity and training data. In their reported tests, it cuts MAPE by 1.5–3x across interpolation and extrapolation. Paired with a sparse low-compute grid, it predicts full-grid behavior with about 10x less compute than uniform sweeps. HF Daily Papers' note
The authors propose “Skaling,” which adds a single interaction exponent coupling capacity and training data. In their reported tests, it cuts MAPE by 1.5–3x across interpolation and extrapolation. Paired with a sparse low-compute grid, it predicts full-grid behavior with about 10x less compute than uniform sweeps. HF Daily Papers' note
score 6