Protoreasoning in Tiny Transformers
Tiny models improved out-of-distribution bracket reasoning when given meaningful intermediate traces.
The paper tests roughly 1M-parameter transformers on Dyck-language tasks, using bracket nesting as a controlled stand-in for reasoning. Valle and Reid call the trace-based method “protoreasoning,” a simpler analogue of chain-of-thought. Ablations say the gains came from the trace content, not just from adding more tokens. ArXiv · AI/CL/LG's note
The paper tests roughly 1M-parameter transformers on Dyck-language tasks, using bracket nesting as a controlled stand-in for reasoning. Valle and Reid call the trace-based method “protoreasoning,” a simpler analogue of chain-of-thought. Ablations say the gains came from the trace content, not just from adding more tokens. ArXiv · AI/CL/LG's note
score 5