Modular TTT: Rethinking Test-Time Training as Composable Modules
The paper turns test-time training variants into swappable graph components, then uses that setup to test which pieces actually matter.
Modular TTT exposes the fast-weight network, loss, learning rate, weight decay, and normalization as explicit design choices. In the authors’ ablations, small learning-rate initialization, weight decay, and a single-layer nonlinearity helped, while MSE and inner-product losses were about even. Deeper fast-weight networks and normalization tended to hurt by producing overly large activations; residuals and gating showed little measurable gain. The best variant was trained at 410M and 1.45B parameters on 100B tokens, with loss and benchmark results comparable to Gated DeltaNet. HF Daily Papers' note
Modular TTT exposes the fast-weight network, loss, learning rate, weight decay, and normalization as explicit design choices. In the authors’ ablations, small learning-rate initialization, weight decay, and a single-layer nonlinearity helped, while MSE and inner-product losses were about even. Deeper fast-weight networks and normalization tended to hurt by producing overly large activations; residuals and gating showed little measurable gain. The best variant was trained at 410M and 1.45B parameters on 100B tokens, with loss and benchmark results comparable to Gated DeltaNet. HF Daily Papers' note
score 5