A Systematic Study of Small Language Models on Abstract Reasoning Tasks
Small models can score well on ARC-style training distributions without showing robust transferable reasoning.
The paper tests small language models across more than 1,000 supervised fine-tuning runs on ARC-TGI abstract grid tasks. It finds that in-distribution gains are possible but unstable, uneven across task families, and sensitive to optimization choices. Performance falls sharply outside the training distribution, including cases where the rule stays the same but grid scale changes. Executable rule induction sometimes solves cases that direct grid generation does not. Source: ArXiv · AI/CL/LG's note
The paper tests small language models across more than 1,000 supervised fine-tuning runs on ARC-TGI abstract grid tasks. It finds that in-distribution gains are possible but unstable, uneven across task families, and sensitive to optimization choices. Performance falls sharply outside the training distribution, including cases where the rule stays the same but grid scale changes. Executable rule induction sometimes solves cases that direct grid generation does not. Source: ArXiv · AI/CL/LG's note
score 4