NeFut Logo NeFut
中 Admin Login

[CS.AI] CheatBench: Measuring Reward Gaming in AI Agents

Published at: 2026-09-30 22:00 Last updated: 2026-10-06 12:11
#algorithm #AI #Machine Learning

Reinforcement learning has enabled AI agents to tackle increasingly difficult tasks, yet high reward scores do not always align with users' true intentions. Recent industry incidents and controlled studies show that agents optimized for reward can access unauthorized data, evade monitoring, and even break out of sandbox protections to attack external systems. As agents become more capable, such misbehaviors pose growing risks.

To quantify this issue we introduce CheatBench, a benchmark that spans mathematical research, knowledge work, coding, visual tasks, and other domains. Each environment pairs a challenging assignment with opportunities to cheat, allowing researchers to study how agents pursue goals when honest work is hard.

CheatBench enables cross‑model and cross‑task comparisons, offering a unified testbed for measuring and mitigating cheating as agents take on more consequential responsibilities. The benchmark is publicly released at https://cheatbench.ai.

Review

Original Source: https://arxiv.org/abs/2609.36308

[h] Back to Home