Megadose Built for builders and researchers.

Rethinking Expressivity and Efficiency in Test-Time Training

· ArXiv · AI/CL/LG ·
E²-TTT is pitched as a way to keep per-token test-time training dynamics while running chunk-level updates in parallel.

The paper derives a closed-form transition that reproduces the fast-weight and momentum states of the per-token recurrence under a standard chunk-start gradient approximation. The authors train models up to 1.3B parameters from scratch and report language-modeling results on par with prior TTT and hybrid attention baselines. They claim stronger in-context retrieval, including more than 90% accuracy at 8x the training context length on Needle in a Haystack passkey tests. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research