From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
The paper proposes scoring pentesting agents by validated vulnerability discovery, not by closed-form benchmark wins.
The authors argue that CTF-style goals, exploit reproduction, and trajectory matching miss the open-ended choices real targets require. Their protocol uses structured ground truth, LLM-based semantic matching, ambiguity-aware scoring, repeated runs for stochastic agents, and efficiency metrics. They also say they are releasing expert-annotated ground truth and code for the evaluation setup. HF Daily Papers' note
The authors argue that CTF-style goals, exploit reproduction, and trajectory matching miss the open-ended choices real targets require. Their protocol uses structured ground truth, LLM-based semantic matching, ambiguity-aware scoring, repeated runs for stochastic agents, and efficiency metrics. They also say they are releasing expert-annotated ground truth and code for the evaluation setup. HF Daily Papers' note
score 5