Megadose Built for builders and researchers.

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

· HF Daily Papers ·
A 4-bit 60B model recovered by QAH matched or beat its bfloat16 source on 7 of 9 benchmarks.

The paper says Quantization-Aware Healing distills the quantized student directly from the original uncompressed model, instead of training against hard labels after compression. In their GPT-OSS 120B to 60B to MXFP4 pipeline, the result uses about four times less weight memory and half the teacher’s parameters. Compared with their QAT baseline, QAH hit a similar peak about seven times faster and stayed stable with continued training. The authors release the model as Hypernova-60B and include deployment notes, including a reported quality gap between distributed-training backends. HF Daily Papers' note

score 4

Categories: Research