AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
AgentCompass separates agent evaluation into benchmark, harness, and environment layers so tests can be recombined without rebuilding the pipeline.
The paper presents it as an open-source, lightweight infrastructure for evaluating LLM-based agents. It adds a fault-tolerant asynchronous runtime and trajectory analysis tools meant to expose failure modes such as reward-hacking. The authors say it supports more than 20 benchmarks across five capability dimensions. HF Daily Papers' note
The paper presents it as an open-source, lightweight infrastructure for evaluating LLM-based agents. It adds a fault-tolerant asynchronous runtime and trajectory analysis tools meant to expose failure modes such as reward-hacking. The authors say it supports more than 20 benchmarks across five capability dimensions. HF Daily Papers' note
score 5