Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
A rank-8 LoRA at one early layer pushed frozen transformers much farther through in-context reference chains.
The paper says 13 pretrained base models reliably followed only 1.4 to 3.6 lines of chain computation. With the small task-trained LoRA, Qwen3-8B rose from 15.5% to 99% exact accuracy on 24-line chains, and longer training reached 50 lines. The authors describe a relay across middle layers, where chain identity passes forward and frozen heads read progressively farther. Removing parent-line attention breaks that relay, and task-specific LoRAs also helped MuSiQue. HF Daily Papers' note
The paper says 13 pretrained base models reliably followed only 1.4 to 3.6 lines of chain computation. With the small task-trained LoRA, Qwen3-8B rose from 15.5% to 99% exact accuracy on 24-line chains, and longer training reached 50 lines. The authors describe a relay across middle layers, where chain identity passes forward and frozen heads read progressively farther. Removing parent-line attention breaks that relay, and task-specific LoRAs also helped MuSiQue. HF Daily Papers' note
score 6