Megadose AI progress, ranked and analyzed.

Full-bandwidth transformer

· HF Daily Papers ·
The paper adds latent feedback so a model can feed its last hidden state into the next decoding step, not just the sampled token.

That feedback is fused with the token embedding through a gated linear unit while keeping the standard transformer setup, KV cache, and language-modeling objective. The authors train 1B-parameter models up to 400B tokens using a scheduled multi-pass objective to preserve parallel teacher forcing. They report gains in validation loss, 5-shot evaluation, math and coding generation, and instruction tuning, with negligible per-token decoding overhead. The models match or approach standard transformers trained on about 1.5x more tokens and can produce shorter reasoning traces at equal or better accuracy. HF Daily Papers' note

score 5

Categories: Research