Megadose Built for builders and researchers.

Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

· HF Daily Papers ·
Updating on a model’s own generations made long-horizon test-time training worse on independent human text.

The paper reports the effect across 128K-token streams and multiple model settings, including TTT-E2E configurations and Adam updates to Qwen3-4B weights. Matched tests suggest the failure is not adaptation itself: similar updates can help on real text, but self-generated feedback creates a closed loop that stores damage. A frozen-generator comparison removed most of the loss in smaller configurations, while replay and one-update tests isolated the conflict between fitting the generated source and predicting new real text. The authors propose “Settlement,” which checks candidate updates on independent real text before committing them, preserving real-text adaptation in their reported endpoints. HF Daily Papers' note

score 6

Categories: Research