Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
SCALE freezes the pretrained model and SFT delta, then learns gates that can suppress, reverse, or extrapolate SFT features.
The paper argues that existing token-reweighting methods can only dampen or amplify supervised updates because their token coefficients stay nonnegative. Its proposed method, SCALE, uses local entropy to learn bounded token- and module-specific gates over already learned SFT residuals. In the reported Qwen math-model tests, SCALE beats the strongest baselines on mathematical reasoning averages while staying competitive on retention benchmarks. It also posts the best average code-generation performance across HumanEval, HumanEval+, and MBPP for all three tested models. HF Daily Papers' note
The paper argues that existing token-reweighting methods can only dampen or amplify supervised updates because their token coefficients stay nonnegative. Its proposed method, SCALE, uses local entropy to learn bounded token- and module-specific gates over already learned SFT residuals. In the reported Qwen math-model tests, SCALE beats the strongest baselines on mathematical reasoning averages while staying competitive on retention benchmarks. It also posts the best average code-generation performance across HumanEval, HumanEval+, and MBPP for all three tested models. HF Daily Papers' note
score 4