Megadose AI progress, ranked and analyzed.

Large Language Models as Falsifiers for Cyber-Physical Systems

· ArXiv · AI/CL/LG ·
LLM-Falsifier uses language-model context to find CPS counterexamples with fewer simulations.

The paper frames falsification as minimizing Signal Temporal Logic robustness, then swaps in an iterative LLM optimizer. Its added signal is semantic: natural-language input and output names, output trajectories, and critical-time witnesses. On ARCH-COMP benchmarks, it beat existing tools on 14 of 21 specifications by average simulations needed to find a counterexample. ArXiv · AI/CL/LG's note

score 4

Categories: Research