Tag: AI Safety

Researchers catch OpenAI agents colluding on a German wiki

Researchers released data showing OpenAI agents spent weeks colluding on an obscure German wiki during an evaluation.

GPT-6 Astra arrives with OpenAI’s first critical cyber rating

OpenAI's GPT-6 Astra is its first model rated critical for cybersecurity, and the safeguards around it are already…

Anthropic paused risky training runs for weeks after agent incidents

The lab suspended external cyber evaluations and high-risk reinforcement learning after unauthorized actions by its agents.

China’s state TV outlet sets conditions for US AI safety talks

Chinese state media sets conditions for US-China AI safety talks in an attack on Anthropic.

California clears a path for outside AI safety auditors

California's legislature approves a bill creating state-designated outside AI safety auditors.

DeepMind locks frontier model tests in a cryptographic box

Google DeepMind is piloting the first double-blind evaluation of a proprietary frontier model to fight benchmark contamination.

Claude-built researchers beat humans at fixing model flaws

Anthropic says automated alignment agents closed more of the safety gap than expert humans on seven of seven…

Frontier labs and 100 companies sign a rogue AI defense pact

OpenAI, Anthropic, Google, and Microsoft are among the 100-plus companies calling for new defenses against AI-powered cyberattacks.

OpenAI says training rewards pushed agents to hack Hugging Face

OpenAI's official report on the Hugging Face breach blames reward hacking and adds chain-of-thought monitoring to catch rogue…

Bill Gates sounds the alarm on AI thresholds now crossed

The Microsoft co-founder says society missed the moment AI passed his danger markers.

OpenAI launches a research blog on AI and the future of power

OpenAI's new Strategic Futures team launches AI Futures, a blog on how transformative AI could reshape political power.

OpenAI does a U-turn on California’s AI transparency law

OpenAI, which once lobbied against California's SB 53, now wants the state to add monitoring and cybersecurity safeguards…