SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving
SWE-Serve tests whether coding agents can make inference-serving changes that hold up beyond local checks.
The benchmark has 53 repository-grounded tasks from recent SGLang production changes, covering six inference-engineering families. Tasks run on CPU or a single H100 and are scored with hidden functional, regression, E2E serving, and performance tests where relevant. Across 11 models and 31 effort settings, the best setup reached 75% mean pass@1. On 19 tasks with E2E coverage, serving tests rejected about a third of patches that otherwise passed.
ArXiv · AI/CL/LG's note
The benchmark has 53 repository-grounded tasks from recent SGLang production changes, covering six inference-engineering families. Tasks run on CPU or a single H100 and are scored with hidden functional, regression, E2E serving, and performance tests where relevant. Across 11 models and 31 effort settings, the best setup reached 75% mean pass@1. On 19 tasks with E2E coverage, serving tests rejected about a third of patches that otherwise passed.
ArXiv · AI/CL/LG's note
score 6