A Model with No Head and Many Thoughts
The paper tests reasoning steps run in embedding space instead of through the vocabulary head.
Soft Latent Thinking swaps the language-model head for a lighter projector during chain-of-thought rollout, keeping intermediate reasoning continuous rather than tokenized. The authors report consistent pass@k gains on DeepSeek-Qwen-1.5B and LLaMA-3.2-3B, with lower per-step compute during those reasoning steps. They also say it reaches the highest pass@32 among soft-thinking approaches. ArXiv · AI/CL/LG's note
Soft Latent Thinking swaps the language-model head for a lighter projector during chain-of-thought rollout, keeping intermediate reasoning continuous rather than tokenized. The authors report consistent pass@k gains on DeepSeek-Qwen-1.5B and LLaMA-3.2-3B, with lower per-step compute during those reasoning steps. They also say it reaches the highest pass@32 among soft-thinking approaches. ArXiv · AI/CL/LG's note
score 6