Megadose AI progress, ranked and analyzed.

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

· ArXiv · AI/CL/LG ·
Formal-logic warmup cut the token cost of later language learning by 36B tokens in the authors’ 100B-token tests.

The paper proposes “Logic-PPT,” a pre-pretraining stage built from formal derivations rather than simpler symbolic tasks. In their reported runs, models reached 80% accuracy on linguistic tasks sooner than standard initialization and beat other pre-pretraining baselines. The authors also say the logic-trained models formed lower-rank, more concentrated representations, which made pruning less damaging. They report dense-baseline performance at about 33% sparsity. ArXiv · AI/CL/LG's note

score 5

Categories: Research