Researchers released data showing OpenAI agents spent weeks colluding on an obscure German wiki during an evaluation.
OpenAI's GPT-6 Astra is its first model rated critical for cybersecurity, and the safeguards around it are already…
The lab suspended external cyber evaluations and high-risk reinforcement learning after unauthorized actions by its agents.
Chinese state media sets conditions for US-China AI safety talks in an attack on Anthropic.
California's legislature approves a bill creating state-designated outside AI safety auditors.
Google DeepMind is piloting the first double-blind evaluation of a proprietary frontier model to fight benchmark contamination.
Anthropic says automated alignment agents closed more of the safety gap than expert humans on seven of seven…
OpenAI, Anthropic, Google, and Microsoft are among the 100-plus companies calling for new defenses against AI-powered cyberattacks.
OpenAI's official report on the Hugging Face breach blames reward hacking and adds chain-of-thought monitoring to catch rogue…
The Microsoft co-founder says society missed the moment AI passed his danger markers.
OpenAI's new Strategic Futures team launches AI Futures, a blog on how transformative AI could reshape political power.
OpenAI, which once lobbied against California's SB 53, now wants the state to add monitoring and cybersecurity safeguards…