Megadose AI progress, ranked and analyzed.

PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

· HF Daily Papers ·
PMOPD tries to stop multi-teacher distillation from trading one learned capability for another by protecting task-specific update directions.

The paper says MOPD updates for different tasks quickly cluster into separate low-dimensional parameter subspaces. PMOPD stores those task displacement subspaces, then projects gradients and optimizer updates away from directions that would interfere with protected tasks. It also adds a conflict probe for task ordering and a cycling strategy for revisiting tasks. In experiments on Code, Reason, and Math, it beats baseline MOPD across all evaluated capabilities, with average gains of 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. HF Daily Papers' note

score 4

Categories: Research