An End-to-End Agent Auditing Engine
A²E tests agent harnesses as full systems, not just by answer correctness.
The paper introduces an Agent Auditing Engine that uses an Agent Task Protocol to plug evaluation tasks into different harnesses. Its monitor automatically records standardized execution traces, then scores harness behavior across efficiency, tool use, planning, and recovery. The experiments found large variation across model-harness pairings, with no one combination winning across every task. ArXiv · AI/CL/LG's note
The paper introduces an Agent Auditing Engine that uses an Agent Task Protocol to plug evaluation tasks into different harnesses. Its monitor automatically records standardized execution traces, then scores harness behavior across efficiency, tool use, planning, and recovery. The experiments found large variation across model-harness pairings, with no one combination winning across every task. ArXiv · AI/CL/LG's note
score 4