Megadose AI progress, ranked and analyzed.

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

· HF Daily Papers ·
A 4B fine-tuned Qwen student beat GPT-5.6 on the benchmark’s deterministic transit-kiosk score.

MetroLLM-Bench tests language models as the policy layer for transit kiosks across 955 cases and six real metro systems. The tasks require structured tool calls and a machine-renderable final state covering outcomes, fares, and kiosk actions. On the held-out split, the 4B Qwen 3.5 PEFT model scored 91.3 on Tier 1, ahead of the two GPT-5.6 results listed at 90.6 and 90.0. The paper says larger 9B and 27B students did not improve Tier 1 at this training scale. HF Daily Papers' note

score 5

Categories: Research