When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis
The paper finds merged specialist agents can match joint RL on AppWorld task completion.
McClendon trains two Qwen3-8B specialists with LOOP, merges them with TIES and RAM+, and compares them with a jointly trained model on the same data. The merge variants are statistically indistinguishable from joint training on task-goal completion. The paper argues this happens because the specialists’ task vectors are near-orthogonal despite substantial support overlap, making the tested merge methods behave close to uniform averaging. Source: ArXiv · AI/CL/LG's note.
McClendon trains two Qwen3-8B specialists with LOOP, merges them with TIES and RAM+, and compares them with a jointly trained model on the same data. The merge variants are statistically indistinguishable from joint training on task-goal completion. The paper argues this happens because the specialists’ task vectors are near-orthogonal despite substantial support overlap, making the tested merge methods behave close to uniform averaging. Source: ArXiv · AI/CL/LG's note.
score 5