Does Learning Protein Folding Generalize to Broader Reasoning?
Protein-folding supervision gave a general LLM measurable gains beyond protein tasks.
The paper introduces FoldingCorpus and Fold2Reason, a post-training method built from protein-derived question-answer data and shared 3D geometry signals. On FoldBench, it reports structure prediction scores 2.7 to 3.5 times higher than Qwen3.5-9B. The model also improved across 10 spatial, graph, scientific, and general reasoning benchmarks, lifting macro-average accuracy from 45.09% to 48.33%. Controls using random, synthetic, or shuffled structure produced smaller or negative gains. HF Daily Papers' note
The paper introduces FoldingCorpus and Fold2Reason, a post-training method built from protein-derived question-answer data and shared 3D geometry signals. On FoldBench, it reports structure prediction scores 2.7 to 3.5 times higher than Qwen3.5-9B. The model also improved across 10 spatial, graph, scientific, and general reasoning benchmarks, lifting macro-average accuracy from 45.09% to 48.33%. Controls using random, synthetic, or shuffled structure produced smaller or negative gains. HF Daily Papers' note
score 5