PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
PlannerForge reports end-to-end LLM-agent testing for autonomous-driving motion planners, with open models close to commercial APIs on most tasks.
The paper frames scenario-based ADS testing as a fragmented pipeline and proposes one LLM-agent framework spanning generation, selection, modification, routing, testing, enhancement, and benchmarking. In its evaluation, best task scores range from 0.88 to 1.00 across 10 LLMs and five prompt conditions. The authors report stronger executable scenario generation than Scenario Factory 2.0, better rank-1 selection than BM25, and more physically valid edits than From-Words-to-Collisions. At N=400, cost tuning raised planner success from 50.4% to 70.2% and reduced collisions from 19.0% to 8.4%, without domain-specific fine-tuning. ArXiv · AI/CL/LG's note
The paper frames scenario-based ADS testing as a fragmented pipeline and proposes one LLM-agent framework spanning generation, selection, modification, routing, testing, enhancement, and benchmarking. In its evaluation, best task scores range from 0.88 to 1.00 across 10 LLMs and five prompt conditions. The authors report stronger executable scenario generation than Scenario Factory 2.0, better rank-1 selection than BM25, and more physically valid edits than From-Words-to-Collisions. At N=400, cost tuning raised planner success from 50.4% to 70.2% and reduced collisions from 19.0% to 8.4%, without domain-specific fine-tuning. ArXiv · AI/CL/LG's note
score 4