APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
A deterministic passport check cut forbidden-recipient payments from 140 cases to zero in the paper’s replayed attack set.
APort Vault tests payment authorization for tool-using agents using 4,371 human-written attacks from a public capture-the-flag event. The benchmark ran 225,964 evaluations across 14 models, five policy configurations, and two replay tracks. Model-only agents still attempted transfers outside the permitted passport recipients at Levels 2 to 4; with the Open Agent Passport pre-action layer, the paper reports none across 69,297 evaluations. The layer did not simply block payments: 25,370 payments executed behind it, while 187 transfer calls were denied by policy. Source: HF Daily Papers' note
APort Vault tests payment authorization for tool-using agents using 4,371 human-written attacks from a public capture-the-flag event. The benchmark ran 225,964 evaluations across 14 models, five policy configurations, and two replay tracks. Model-only agents still attempted transfers outside the permitted passport recipients at Levels 2 to 4; with the Open Agent Passport pre-action layer, the paper reports none across 69,297 evaluations. The layer did not simply block payments: 25,370 payments executed behind it, while 187 transfer calls were denied by policy. Source: HF Daily Papers' note
score 5