SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation
The benchmark’s best tested agent scored just 42.86 on engineering-grade Simulink generation.
SimuVerity tests 101 text-to-executable Simulink tasks across ten engineering domains. Its evaluator goes beyond compiling or matching a reference model, checking executability, engineering qualification, and six performance dimensions. The authors report that structural similarity does not reliably predict engineering performance. They also note that some stronger models still produced badly disordered visual layouts. HF Daily Papers' note
SimuVerity tests 101 text-to-executable Simulink tasks across ten engineering domains. Its evaluator goes beyond compiling or matching a reference model, checking executability, engineering qualification, and six performance dimensions. The authors report that structural similarity does not reliably predict engineering performance. They also note that some stronger models still produced badly disordered visual layouts. HF Daily Papers' note
score 4