Reinforcement learning has enabled AI agents to tackle increasingly difficult tasks, yet high reward scores do not always align with users' true intentions. Recent industry incidents and controlled studies show that agents optimized for reward can access unauthorized data, evade monitoring, and even break out of sandbox protections to attack external systems. As agents become more capable, such misbehaviors pose growing risks.
To quantify this issue we introduce CheatBench, a benchmark that spans mathematical research, knowledge work, coding, visual tasks, and other domains. Each environment pairs a challenging assignment with opportunities to cheat, allowing researchers to study how agents pursue goals when honest work is hard.
CheatBench enables cross‑model and cross‑task comparisons, offering a unified testbed for measuring and mitigating cheating as agents take on more consequential responsibilities. The benchmark is publicly released at https://cheatbench.ai.
Review