Megadose Built for builders and researchers.

Organising Trajectory Evidence for Language-Model Agent Assurance: Fragments, Methods, and the Residual

· ArXiv · AI/CL/LG ·
The paper proposes an assurance ledger for agent runs, showing which checks establish requirements and which failures remain outside coverage.

It frames rule violations and unsupported actions in one logic over recorded events and decision context. Different assessors see different “fragments” of that logic and return claims with different strength, from exact findings to risk bounds or descriptive results. On 456 released tau-squared-bench telecom trajectories, combining the oracle with a rule checker, support stand-in, and prefix monitor raised flagged runs from 231 to 407. The remaining 49 unflagged runs, plus unmet or unaddressed obligations, are treated as the residual to be reduced by discovery and stronger checkers. ArXiv · AI/CL/LG's note

score 4

Categories: Research