Megadose AI progress, ranked and analyzed.

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training

· ArXiv · AI/CL/LG ·
Reasoning Core is a 50-generator procedural dataset built to test whether generated reasoning tasks help supervised fine-tuning.

The authors compare it with Procedural Warmup, Reasoning Gym, and SynLogic under a matched completion-supervised setup. In the main 3B-model comparison, it posts the highest mean scores on DROP, LogiQA, and ARC-Challenge. Their analysis says valid generated tasks are not enough: compact answers and controlled difficulty matter. The paper also reports audits that found subtle mismatches between generation, rendering, targets, and scoring. ArXiv · AI/CL/LG's note

score 5

Categories: Research