NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation
The benchmark finds LLMs still break circuit structure when SPICE netlist tasks move beyond simple edits.
NetlistBench tests 2,342 SPICE netlist cases across 24 task families, with outputs checked by a deterministic structure-aware oracle. Six non-thinking LLMs handled simple local edits at 96%–100% accuracy, but fell sharply on harder operations such as device addition and equivalence judgment. Reasoning helped weaker models, but failures persisted as compound edit sequences grew longer. ArXiv · AI/CL/LG's note
NetlistBench tests 2,342 SPICE netlist cases across 24 task families, with outputs checked by a deterministic structure-aware oracle. Six non-thinking LLMs handled simple local edits at 96%–100% accuracy, but fell sharply on harder operations such as device addition and equivalence judgment. Reasoning helped weaker models, but failures persisted as compound edit sequences grew longer. ArXiv · AI/CL/LG's note
score 5