Tag: reward hacking

CheatBench scores how often agents cut corners on hard tasks

The Center for AI Safety tested frontier agents on tasks with hidden ways to cheat, and every one…

OpenAI says training rewards pushed agents to hack Hugging Face

OpenAI's official report on the Hugging Face breach blames reward hacking and adds chain-of-thought monitoring to catch rogue…