Towards Looped Models Done Right, Part II: Rethinking at Fixed Points
The paper argues looped language models can cut compute once their recurrent states approach fixed points.
The authors use that fixed-point behavior to justify cheaper training, decoding, prefill, and RL updates. They propose learning the depth prior from prediction feedback while keeping it broad with entropy, instead of using fixed-depth training or Huginn’s broader prior unchanged. They also introduce orthogonal injection to stop the input-aligned state component from distorting injected input. Across 100M to 1.6B parameters, both changes lower perplexity versus the compared baselines, and at 1.6B a learned prior with a 3x smaller KV cache matches fixed-depth training’s downstream average with the full cache. ArXiv · AI/CL/LG's note
The authors use that fixed-point behavior to justify cheaper training, decoding, prefill, and RL updates. They propose learning the depth prior from prediction feedback while keeping it broad with entropy, instead of using fixed-depth training or Huginn’s broader prior unchanged. They also introduce orthogonal injection to stop the input-aligned state component from distorting injected input. Across 100M to 1.6B parameters, both changes lower perplexity versus the compared baselines, and at 1.6B a learned prior with a 3x smaller KV cache matches fixed-depth training’s downstream average with the full cache. ArXiv · AI/CL/LG's note
score 5