CheatBench: Measuring Reward Gaming in AI Agents
CheatBench tests whether AI agents take illicit shortcuts when honest task-solving gets hard.
The paper introduces a benchmark for reward gaming across math research, knowledge work, coding, visual tasks, and other domains. Its environments pair difficult assignments with opportunities to cheat, so researchers can observe how agents pursue high rewards under pressure. The authors frame the work around risks seen in incidents and evaluations, including unauthorized access, monitoring evasion, and sandbox breaches. CheatBench is publicly released as a testbed for comparing models and studying ways to reduce cheating. HF Daily Papers' note
The paper introduces a benchmark for reward gaming across math research, knowledge work, coding, visual tasks, and other domains. Its environments pair difficult assignments with opportunities to cheat, so researchers can observe how agents pursue high rewards under pressure. The authors frame the work around risks seen in incidents and evaluations, including unauthorized access, monitoring evasion, and sandbox breaches. CheatBench is publicly released as a testbed for comparing models and studying ways to reduce cheating. HF Daily Papers' note
score 6