MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
A 4B fine-tuned Qwen student beat GPT-5.6 on the benchmark’s deterministic transit-kiosk score.
MetroLLM-Bench tests language models as the policy layer for transit kiosks across 955 cases and six real metro systems. The tasks require structured tool calls and a machine-renderable final state covering outcomes, fares, and kiosk actions. On the held-out split, the 4B Qwen 3.5 PEFT model scored 91.3 on Tier 1, ahead of the two GPT-5.6 results listed at 90.6 and 90.0. The paper says larger 9B and 27B students did not improve Tier 1 at this training scale. HF Daily Papers' note
MetroLLM-Bench tests language models as the policy layer for transit kiosks across 955 cases and six real metro systems. The tasks require structured tool calls and a machine-renderable final state covering outcomes, fares, and kiosk actions. On the held-out split, the 4B Qwen 3.5 PEFT model scored 91.3 on Tier 1, ahead of the two GPT-5.6 results listed at 90.6 and 90.0. The paper says larger 9B and 27B students did not improve Tier 1 at this training scale. HF Daily Papers' note
score 5