On the Threat Model of Weird Generalization and Emergent Misalignment
The paper argues weird generalization looks fragile enough to be an adversarially engineered risk, not a routine fine-tuning hazard.
The authors test three open-weight models across four datasets to see which data features make narrow fine-tuning produce broad behavior shifts. Dataset composition and language mattered more than dataset size. Weird generalization was stronger when the fine-tuning data was familiar from pretraining than when it was novel. The measurement itself also shifted depending on the small evaluation question set used. ArXiv · AI/CL/LG's note
The authors test three open-weight models across four datasets to see which data features make narrow fine-tuning produce broad behavior shifts. Dataset composition and language mattered more than dataset size. Weird generalization was stronger when the fine-tuning data was familiar from pretraining than when it was novel. The measurement itself also shifted depending on the small evaluation question set used. ArXiv · AI/CL/LG's note
score 4