NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents
NeutronGym tests LLM agents by making their instrument designs run through physics, not an LLM judge.
The paper introduces an executable neutron instrument design environment where agents set parameters, McStas ray-traces the result, and a grading ladder scores syntax, runtime, structure, and science. In McStasBench, seven models solved at most 7 of 16 published-instrument tasks, with none retrieving a reference or meeting the improvement target. Reinforcement learning improved Qwen3-8B sharply on held-out procedural instances, but the paper says the gain depended on partial-credit rewards. The authors also report releasing probes that exposed four task designs solvable by no-model baselines. ArXiv · AI/CL/LG's note
The paper introduces an executable neutron instrument design environment where agents set parameters, McStas ray-traces the result, and a grading ladder scores syntax, runtime, structure, and science. In McStasBench, seven models solved at most 7 of 16 published-instrument tasks, with none retrieving a reference or meeting the improvement target. Reinforcement learning improved Qwen3-8B sharply on held-out procedural instances, but the paper says the gain depended on partial-credit rewards. The authors also report releasing probes that exposed four task designs solvable by no-model baselines. ArXiv · AI/CL/LG's note
score 5