Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation
A 34M-parameter correction layer fixed about half the frozen model’s tested errors without measurable benchmark damage.
The paper’s CRN v2 module sits on top of a frozen Gemma 4 E2B model and trains only the correction layer. On a 60-question CEHRI exam, it corrected 53.3% of base-model errors, or 43.3% on a reworded variant. The tested MMLU, BoolQ, and small car-wash benchmarks showed no degradation, while a matched LoRA baseline corrected more errors but lost 30–75% capability on those checks. The author frames this as evidence for the frozen-base, logit-correction, KL-anchored design principle, not a new architecture claim. HF Daily Papers' note
The paper’s CRN v2 module sits on top of a frozen Gemma 4 E2B model and trains only the correction layer. On a 60-question CEHRI exam, it corrected 53.3% of base-model errors, or 43.3% on a reworded variant. The tested MMLU, BoolQ, and small car-wash benchmarks showed no degradation, while a matched LoRA baseline corrected more errors but lost 30–75% capability on those checks. The author frames this as evidence for the frozen-base, logit-correction, KL-anchored design principle, not a new architecture claim. HF Daily Papers' note
score 4