Megadose AI progress, ranked and analyzed.

Fusion Training for Mathematical Generalization in Large Language Models

· ArXiv · AI/CL/LG ·
More non-thinking supervision made the model’s thinking mode worse on math tasks.

The paper studies Thinking Mode Fusion, where one model handles both concise answers and longer reasoning. Its benchmark varies the ratio of thinking to non-thinking data and tests three training schedules. The authors report an asymmetric trade-off: adding more non-thinking data reduces thinking-mode accuracy, while schedules change how severe that tension is. Code and data are released as Fusion Bench. ArXiv · AI/CL/LG's note

score 5

Categories: Research