Megadose AI progress, ranked and analyzed.

Thought without systematicity? Evaluating reasoning models on rule induction tasks

· HF Daily Papers ·
Reasoning models can solve a rule induction task and still break on an equivalent version of it.

Schug and Lake test whether current reasoning models show systematicity: stable performance across structurally matched variants of the same task. They adapt cognitive-science rule induction tasks and create equivalent variations through recombination and substitution. The paper reports that models often fail those variants even after succeeding on the original form, making their apparent reasoning ability hard to generalize beyond the exact evaluation context. HF Daily Papers' note

score 5

Categories: Research