AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation
The paper argues that RAG releases should be gated by measured evidence, not clean-looking judge output.
AGO uses a four-state release model, layered checks, probabilistic regression gating, and mandatory validation of the LLM judge. In the public RAGBench evaluation, gpt-4.1-nano followed the protocol but detected bad answers only barely above chance. gpt-4o performed better, though its results varied sharply by domain. In simulated regression cases, the stricter profile cut unsafe promotions versus a naive gate.
HF Daily Papers' note
AGO uses a four-state release model, layered checks, probabilistic regression gating, and mandatory validation of the LLM judge. In the public RAGBench evaluation, gpt-4.1-nano followed the protocol but detected bad answers only barely above chance. gpt-4o performed better, though its results varied sharply by domain. In simulated regression cases, the stricter profile cut unsafe promotions versus a naive gate.
HF Daily Papers' note
score 4