Allspark: Weak to Strong Transfer via Alternating Chain of Thought
Allspark trains a weak model to guide stronger models through alternating reasoning, without using strong-model rollouts in training.
The paper pairs a trainable weak teacher with a frozen same-size model during training, alternating chain-of-thought segments before the frozen model gives the answer. At inference, that frozen partner is replaced by a stronger fixed model, which the weak teacher steers through text. The authors report gains in controlled Qwen tests and larger ARC-AGI-2 experiments, including cross-family transfer to Kimi and Nemotron. They frame the result as a way to reuse one weak teacher across stronger students while studying the accuracy-versus-token cost. HF Daily Papers' note
The paper pairs a trainable weak teacher with a frozen same-size model during training, alternating chain-of-thought segments before the frozen model gives the answer. At inference, that frozen partner is replaced by a stronger fixed model, which the weak teacher steers through text. The authors report gains in controlled Qwen tests and larger ARC-AGI-2 experiments, including cross-family transfer to Kimi and Nemotron. They frame the result as a way to reuse one weak teacher across stronger students while studying the accuracy-versus-token cost. HF Daily Papers' note
score 5