Small Foundation Models of Human Cognition and Behaviour
Models under 1B parameters matched a 70B baseline on held-out participants, but larger models pulled ahead on unfamiliar task structures.
Oh and Gobet trained fourteen models on Psych-101, a 10.7 million-trial dataset from 160 psychology experiments. In-distribution performance hit a narrow ceiling where scale mattered little. Out of distribution, scale mattered more, with bigger models better at generalising to new task structure. Their diagnostics found that masking stimuli and feedback erased most learned information, arguing against choice history alone explaining the results. HF Daily Papers' note
Oh and Gobet trained fourteen models on Psych-101, a 10.7 million-trial dataset from 160 psychology experiments. In-distribution performance hit a narrow ceiling where scale mattered little. Out of distribution, scale mattered more, with bigger models better at generalising to new task structure. Their diagnostics found that masking stimuli and feedback erased most learned information, arguing against choice history alone explaining the results. HF Daily Papers' note
score 5