Megadose AI progress, ranked and analyzed.

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

· HF Daily Papers ·
TraceDance turns messy deployment traces into behavior-specific tests for agent failures.

The system builds targeted benchmarks for undesirable behaviors seen in real agent sessions, rather than relying on fixed suites. Its experiments used 252,557 sessions to produce 107 benchmarks with 4,125 instances, meeting 95.3% of build requests. Human annotators confirmed the requested behavior in 84% of sampled cases, and the automated grader tracked human pass/fail judgments at roughly annotator-level agreement. Nine frontier LLMs averaged a 26.7% pass rate at the evaluated decision points. HF Daily Papers' note

score 5

Categories: Research