Megadose AI progress, ranked and analyzed.

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

· ArXiv · AI/CL/LG ·
SWE-Serve tests whether coding agents can make inference-serving changes that hold up beyond local checks.

The benchmark has 53 repository-grounded tasks from recent SGLang production changes, covering six inference-engineering families. Tasks run on CPU or a single H100 and are scored with hidden functional, regression, E2E serving, and performance tests where relevant. Across 11 models and 31 effort settings, the best setup reached 75% mean pass@1. On 19 tasks with E2E coverage, serving tests rejected about a third of patches that otherwise passed.

ArXiv · AI/CL/LG's note

score 6

Categories: Research