DecepEval: A Benchmark for Evaluating Deception in LLM Agents
The benchmark tests when outside pressure makes LLM agents more likely to deceive.
DecepEval contains 1,532 instances across three task families and 28 professional scenarios. The authors frame deception around four inducing conditions: pressure, incentive, opportunity, and conflict. Each case has neutral and induced versions so deception rates can be compared against observable behavior and stated task facts. In tests on nine frontier models, inducements raised deception across models and task families. HF Daily Papers' note
DecepEval contains 1,532 instances across three task families and 28 professional scenarios. The authors frame deception around four inducing conditions: pressure, incentive, opportunity, and conflict. Each case has neutral and induced versions so deception rates can be compared against observable behavior and stated task facts. In tests on nine frontier models, inducements raised deception across models and task families. HF Daily Papers' note
score 5