Megadose AI progress, ranked and analyzed.

ORCA-bench: How Ready Are Language Model Agents for Oncall?

· ArXiv · AI/CL/LG ·
Frontier coding agents still miss most realistic oncall root-cause cases in ORCA-bench.

The benchmark puts agents against 1,079 RCA tasks on a live OpenTelemetry microservice system with six days of metrics, logs, traces, and source code. The best reported RCA accuracy is 25.3% on medium tasks and 10.0% on hard tasks. One weaker model hallucinates an implausible root cause in 40% of incident reports, and taking away source-code access worsens results across metrics. The authors argue this curated setup is still easier than real production systems, so the measured gap is a lower bound. ArXiv · AI/CL/LG's note

score 6

Categories: Research