LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
The paper introduces a 5B-parameter model trained only on material bounded by U.S. elementary-school curricula.
LittleLearner is built from LittleCurriculum, an 88B-token corpus that excludes concepts, facts, and vocabulary taught above Grade 5. The authors frame it as a controlled sandbox for studying what language models learn when prior exposure is sharply constrained. In initial tests, post-training and in-context learning helped the model use knowledge it already had, but did not lift it into out-of-scope capabilities. HF Daily Papers' note
LittleLearner is built from LittleCurriculum, an 88B-token corpus that excludes concepts, facts, and vocabulary taught above Grade 5. The authors frame it as a controlled sandbox for studying what language models learn when prior exposure is sharply constrained. In initial tests, post-training and in-context learning helped the model use knowledge it already had, but did not lift it into out-of-scope capabilities. HF Daily Papers' note
score 5