Megadose AI progress, ranked and analyzed.

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

· ArXiv · AI/CL/LG ·
A single self-hosted model took over half of internal LLM traffic after targeted post-training on production failures.

The paper says the team consolidated requests from more than 200 internal applications by training separate GRPO experts for instruction following, function calling, and internal task distribution. They merged those experts with two-stage SLERP after finding joint optimization caused reward interference. In non-reasoning mode, the resulting model beat a roughly 7x larger baseline on their in-house Arena, instruction-following, and function-calling tests. It now handles 116 million requests a month at lower serving cost. ArXiv · AI/CL/LG's note

score 5

Categories: Research