Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
The paper says shared LLM endpoints failed the reliability checks needed to treat model judges as stable measurement tools.
In two preregistered audits, repeat rankings landed far below the authors’ fixed reliability thresholds. Same-window repeats reached Spearman 0.400 against a required 0.90, while byte-identical next-day replays reached 0.78 against 0.99. The authors attribute the failure to label effects, tiny candidate gaps relative to noise, and different rankings from identical inputs. Follow-ups found waiting, provider switching, and metric changes did not repair the issue on the tested grid. ArXiv · AI/CL/LG's note
In two preregistered audits, repeat rankings landed far below the authors’ fixed reliability thresholds. Same-window repeats reached Spearman 0.400 against a required 0.90, while byte-identical next-day replays reached 0.78 against 0.99. The authors attribute the failure to label effects, tiny candidate gaps relative to noise, and different rankings from identical inputs. Follow-ups found waiting, provider switching, and metric changes did not repair the issue on the tested grid. ArXiv · AI/CL/LG's note
score 5