Megadose Built for builders and researchers.

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

· HF Daily Papers ·
The paper argues looped language models can cut recurrence costs once their states converge near fixed points.

The authors use that property for cheaper training, decoding, prefill, and RL updates. They propose learning the depth prior from prediction feedback while keeping it broad, and using orthogonal injection to prevent the input component from amplifying or canceling itself. Across 100M to 1.6B parameters, both changes lower perplexity versus the baselines named in the abstract. At 1.6B, their learned prior matches fixed-depth downstream averages with a 3x smaller KV cache. HF Daily Papers' note

score 5

Categories: Research