ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts
The benchmark says final-answer scores miss where runtime agents actually fail.
ClawProBench evaluates declared model-plus-runtime configurations from execution traces, not just answers. It has a 102-scenario full profile and a frozen 68-scenario holdout with JSON output contracts. The paper reports weak alignment between full-profile and holdout rankings, and says correctness-only leaderboards diverge from safety-gated, process-aware scoring. Native-runtime tasks also scored lower than workspace-live tasks in its evaluation. HF Daily Papers' note
ClawProBench evaluates declared model-plus-runtime configurations from execution traces, not just answers. It has a 102-scenario full profile and a frozen 68-scenario holdout with JSON output contracts. The paper reports weak alignment between full-profile and holdout rankings, and says correctness-only leaderboards diverge from safety-gated, process-aware scoring. Native-runtime tasks also scored lower than workspace-live tasks in its evaluation. HF Daily Papers' note
score 5