Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining
The paper tracks which pretraining examples point a model toward its final weights without tying the measurement to any downstream task.
The authors define influence as how much an example’s gradient update reduces distance to the final parameters of a pretraining run. They estimate that from intermediate checkpoints, avoiding retraining. Across 18 Pythia and PolyPythia configurations, literature-related data mattered more early, while STEM data became more aligned later. The paper says that crossover was broadly consistent across configurations. ArXiv · AI/CL/LG's note
The authors define influence as how much an example’s gradient update reduces distance to the final parameters of a pretraining run. They estimate that from intermediate checkpoints, avoiding retraining. Across 18 Pythia and PolyPythia configurations, literature-related data mattered more early, while STEM data became more aligned later. The paper says that crossover was broadly consistent across configurations. ArXiv · AI/CL/LG's note
score 5