The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs
The paper says organizing separate QLoRA training paths by task compatibility beats adding more adapter capacity under the same budget.
The authors target interference and forgetting in PEFT when heterogeneous tasks share one LoRA optimization path. Their framework automatically groups and sequences tasks, then trains independent Quantized Low-Rank Adapters for compatible paths. On TRACE, the automatic multi-policy setup reports the best score, 44.78, with the same trainable capacity. ArXiv · AI/CL/LG's note
The authors target interference and forgetting in PEFT when heterogeneous tasks share one LoRA optimization path. Their framework automatically groups and sequences tasks, then trains independent Quantized Low-Rank Adapters for compatible paths. On TRACE, the automatic multi-policy setup reports the best score, 44.78, with the same trainable capacity. ArXiv · AI/CL/LG's note
score 4