Megadose Built for builders and researchers.

ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts

· HF Daily Papers ·
The benchmark says final-answer scores miss where runtime agents actually fail.

ClawProBench evaluates declared model-plus-runtime configurations from execution traces, not just answers. It has a 102-scenario full profile and a frozen 68-scenario holdout with JSON output contracts. The paper reports weak alignment between full-profile and holdout rankings, and says correctness-only leaderboards diverge from safety-gated, process-aware scoring. Native-runtime tasks also scored lower than workspace-live tasks in its evaluation. HF Daily Papers' note

score 5

Categories: Research