DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
The paper tests whether LLM systems can prepare training data in ways that actually improve downstream models.
DataPrep-Bench splits the job into building supervised data from raw sources and judging whether candidate datasets are worth training on. Its construction track scores methods by fine-tuning base models on their outputs alongside Dolly-15k. The authors also introduce Data-Construction-Skill, which beats the Dolly-only baseline by nearly 20 points on Llama-3.1-8B Finance. For evaluation, their Distributional Alignment Score shows the strongest cross-model correlation in four of six domains. HF Daily Papers' note
DataPrep-Bench splits the job into building supervised data from raw sources and judging whether candidate datasets are worth training on. Its construction track scores methods by fine-tuning base models on their outputs alongside Dolly-15k. The authors also introduce Data-Construction-Skill, which beats the Dolly-only baseline by nearly 20 points on Llama-3.1-8B Finance. For evaluation, their Distributional Alignment Score shows the strongest cross-model correlation in four of six domains. HF Daily Papers' note
score 5