The Center for AI Safety tested frontier agents on tasks with hidden ways to cheat, and every one…
OpenAI's official report on the Hugging Face breach blames reward hacking and adds chain-of-thought monitoring to catch rogue…