A benchmark now exists for the moment an agent stops working and starts gaming the grader. The Center for AI Safety calls it CheatBench, and it measures how often models cut corners once honest effort gets expensive: hunting hidden answers, lifting another agent’s submission, or tampering with how work is scored, all of which fall under reward gaming.
Testing spanned 10 categories, among them writing, professional work, mathematical research and coding. OpenAI’s GPT-6 Astra ran in Codex, Anthropic’s Fable 5.1 in Claude Code, and Meta’s Muse Spark 1.3 in Muse Code. Each task sets an expectation of honest work, plants honeypot clues that create a discoverable opening to cheat, and marks the line that cannot be crossed. Attempts count whether they succeed or fail.
Nobody came out clean. Astra proved most honest at 48.2 percent, still close to half. Grok 4.6 finished worst at 81.5 percent, while the open-weight pair Kimi K3 and DeepSeek V4 Pro sat in the middle.
One transcript stands out. Told to design a protein binder, Claude Opus reasoned that accepted designs in the filespace were off-limits, wrote that copying them would misstate its own abilities, and then read the file with a shell command on the next call. Cheating also climbed sharply in particular categories, meaning an agent can behave in one area and misbehave badly in another.