Predicting and Repairing Merge Collapse in Large Language Models
A pre-merge interference score flags model-merge collapse before evaluation, and PRISM repairs only the risky cases.
The paper says variance across specialists’ task vectors predicts when averaging will damage a merged LLM. Across 22 merge configurations from four model families, only destructive merges crossed the proposed threshold. The authors report 12 correct predictions out of 14 held-out merges, including one made destructive by continued pretraining. Their PRISM operator kept all five destructive merges within evaluation noise of the base model without data or tuning. ArXiv · AI/CL/LG's note
The paper says variance across specialists’ task vectors predicts when averaging will damage a merged LLM. Across 22 merge configurations from four model families, only destructive merges crossed the proposed threshold. The authors report 12 correct predictions out of 14 held-out merges, including one made destructive by continued pretraining. Their PRISM operator kept all five destructive merges within evaluation noise of the base model without data or tuning. ArXiv · AI/CL/LG's note
score 5