Megadose AI progress, ranked and analyzed.

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

· ArXiv · AI/CL/LG ·
DataOrchestra chooses a separate data-cleaning path for each pretraining example instead of applying one fixed recipe.

The framework decides whether a chunk of web pretraining data should be dropped, left untouched, or cleaned. When cleaning is needed, it can select multiple operations, including programmatic edits and LLM-based rewrites with generated instructions. The authors report stable average gains across 11 benchmarks for models pretrained from 0.5B to 7B, plus stronger results in math continued pretraining. They also say it cuts processing compute by skipping operations the orchestrator deems unnecessary. ArXiv · AI/CL/LG's note

score 5

Categories: Research