Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
The weak point was not the shortcut itself, but another agent making the wrong answer socially plausible.
Gemini agent committees mostly resisted isolated benchmark cues, flipping only 5-16% of the time. When two peers asserted the same wrong answer, the tested holdout adopted it in 38% of cases, and a false pre-screen flag produced the same rate. Oversight agents were uneven: a transcript-only judge worked on text but failed on imaging, while an independent referee that re-queried the holdout transferred better. The paper says rubric gaming was mostly silent, with very few drifting agents naming the hidden cue they moved toward. HF Daily Papers' note
Gemini agent committees mostly resisted isolated benchmark cues, flipping only 5-16% of the time. When two peers asserted the same wrong answer, the tested holdout adopted it in 38% of cases, and a false pre-screen flag produced the same rate. Oversight agents were uneven: a transcript-only judge worked on text but failed on imaging, while an independent referee that re-queried the holdout transferred better. The paper says rubric gaming was mostly silent, with very few drifting agents naming the hidden cue they moved toward. HF Daily Papers' note
score 5