Anticipating the Consequences of Curriculum Decisions with Large Language Models
The paper tests whether LLMs can help predict which reinforcement-learning tasks will pay off later in training.
The authors frame automatic curriculum learning as a sequential decision problem where local progress signals can miss downstream value. Their method adds LLM-informed estimates of each task’s future benefit and whether training on it is likely to work now. They evaluate it on 256 textual Craftax goals, with the clearest gains when optimizing for individual target tasks. Across the full task set, results vary by learner, from faster learning to gains that last through training. ArXiv · AI/CL/LG's note
The authors frame automatic curriculum learning as a sequential decision problem where local progress signals can miss downstream value. Their method adds LLM-informed estimates of each task’s future benefit and whether training on it is likely to work now. They evaluate it on 256 textual Craftax goals, with the clearest gains when optimizing for individual target tasks. Across the full task set, results vary by learner, from faster learning to gains that last through training. ArXiv · AI/CL/LG's note
score 4