Megadose AI progress, ranked daily.

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

· HF Daily Papers ·
The benchmark’s best reported root-cause locator still misses the exact decisive step about three-quarters of the time.

LongRCA Bench contains 1,140 failed long-horizon agent trajectories across five domains, with human labels for both the responsible role and the earliest decisive root-cause step. The median trace runs 145 steps, making full manual inspection the problem the paper is trying to measure. The strongest baseline reaches 13.2% exact root-step accuracy. The authors’ training-free RCTA method improves that to 24.1% exact root-step accuracy and 51.1% responsible-role accuracy. HF Daily Papers' note

score 5

Categories: Research