LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles
The paper found no confirmed Autoware weakness; the blockers were mostly build-integration failures.
The authors tested whether LLMs could generate executable dynamic-analysis artifacts for attacker-reachable software sites in an autonomous-driving stack. Across 3,700 generated artifact sets, the strongest model compiled many harnesses only after heavy stubbing, and fewer than half reached fuzzing. All 37 crashes observed came from stubbed code rather than Autoware itself. The paper’s conclusion is that reliable LLM-assisted dynamic threat analysis is currently limited less by finding candidate sites than by wiring generated tests into real, large software builds. ArXiv · AI/CL/LG's note
The authors tested whether LLMs could generate executable dynamic-analysis artifacts for attacker-reachable software sites in an autonomous-driving stack. Across 3,700 generated artifact sets, the strongest model compiled many harnesses only after heavy stubbing, and fewer than half reached fuzzing. All 37 crashes observed came from stubbed code rather than Autoware itself. The paper’s conclusion is that reliable LLM-assisted dynamic threat analysis is currently limited less by finding candidate sites than by wiring generated tests into real, large software builds. ArXiv · AI/CL/LG's note
score 4