Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
Jev is tested as a zero-shot alignment-failure detector and reports a median AUROC of 0.886 with one generic question.
The paper introduces RLCDAlignBench, covering ten failure types across 44 benchmarks and five target models. Its setup separates what Jev is asked from what context it is shown, because many failures depend on reference information outside the model response. The authors say context matters more than question wording, especially when fields encode the benchmark label. They also report that Jev matches reference-scorer agreement with human labels, flags label defects, and is 63x cheaper than LLM-judge scorers. HF Daily Papers' note
The paper introduces RLCDAlignBench, covering ten failure types across 44 benchmarks and five target models. Its setup separates what Jev is asked from what context it is shown, because many failures depend on reference information outside the model response. The authors say context matters more than question wording, especially when fields encode the benchmark label. They also report that Jev matches reference-scorer agreement with human labels, flags label defects, and is 63x cheaper than LLM-judge scorers. HF Daily Papers' note
score 5