The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
The paper argues that LLM math failures often come from missing structural “primitives,” especially at the discovery step, not just weak execution.
The authors introduce a benchmark, HLEI, to test mathematical reasoning across Discovery, Generation, Digestion, and Execution. Their diagnosis says overall answer accuracy hides different capability profiles, and that giving models the right primitives can expose latent execution ability. They identify Discovery as the main bottleneck, and say discovery-limited failures are especially repairable through post-training. They also propose ABS, a primitive-privileged self-distillation method that improves results across model scales and hard benchmarks. HF Daily Papers' note
The authors introduce a benchmark, HLEI, to test mathematical reasoning across Discovery, Generation, Digestion, and Execution. Their diagnosis says overall answer accuracy hides different capability profiles, and that giving models the right primitives can expose latent execution ability. They identify Discovery as the main bottleneck, and say discovery-limited failures are especially repairable through post-training. They also propose ABS, a primitive-privileged self-distillation method that improves results across model scales and hard benchmarks. HF Daily Papers' note
score 5