Megadose AI progress, ranked and analyzed.

PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

· ArXiv · AI/CL/LG ·
PPL-Factory reports stronger math fine-tuning results while using a small slice of the training data.

The paper proposes selecting fine-tuning samples with task-aware perplexity scores and budget-aware criteria. Its authors argue that fixed heuristics like data quality, diversity, or reasoning trace length do not transfer reliably across tasks. On GSM8K, PPL-Factory beats other data-selection methods using 1% of the training set. With 10% of the data, it exceeds full-data fine-tuning accuracy by 0.9 on GSM8K and 4.8 on MATH. ArXiv · AI/CL/LG's note

score 4

Categories: Research