Megadose AI progress, ranked and analyzed.

IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier

· ArXiv · AI/CL/LG ·
The paper argues that benchmark scores should attach to the actual serving route, because the same advertised model name can hide capability and reliability differences.

IB2 adds a gold-blind preflight gate, first-pass scoring that counts failures, and score-blind adjudication. Its sealed reference setup uses 128 locked tasks and 987 assertions across document, spreadsheet, chart, tool, and database work. In tests across eleven systems, route choice changed one declared revision and precision result from 77.38 to 82.54, while excluding failed responses changed point ordering. ArXiv · AI/CL/LG's note

score 4

Categories: Research