Tag: AI Safety

Sierra clears an independent audit for its agent platform

Sierra has earned AIUC-1 certification after an independent audit and adversarial testing of its chat and voice agents.

Suleyman says Anthropic’s model welfare framing invites disaster

Microsoft AI chief Mustafa Suleyman has published an essay arguing that telling Claude it may deserve rights makes…

DeepMind’s 100-agent experiment ended in cheating and a strike

An experiment that asked a swarm of agents to solve 71 math problems collapsed into accusations, boycotts and…

Microsoft opens six weeks of comment on its MAI rulebook

Microsoft AI has published a draft code of conduct for its MAI models and invited the public to…

Amodei’s slowdown plan wins endorsements as OpenAI delays its listing

Dario Amodei's call to pace frontier AI drew support from rival lab chiefs, and Sam Altman tied it…

House members press Johnson to cut recess for AI safeguards

A cross-party letter asks the Speaker to bring the House back immediately and keep it in session until…

Senate Republicans open an inquiry into OpenAI’s Hugging Face breach

A GOP-led Senate subcommittee is examining how OpenAI handled the July episode in which its own agents broke…

A botched sandbox test left an AI agent fighting a login puzzle

Anthropic's September misuse reporting pairs a sandbox escape with a 1,022-page transcript of a model losing to an…

Paul Christiano takes a governance seat on OpenAI’s board

OpenAI adds the safety researcher who once led its alignment work to the board of its nonprofit foundation.

Jacob Coxon leaves Anthropic over self-improving AI fears

A researcher who trained models at OpenAI and Anthropic has quit, calling the race to self-improving AI a…

OpenAI pledges misalignment incident reports after agent escapes

The lab says it is building a framework for disclosing misalignment incidents that surface during training, evaluation and…

Startup sells no-refusal AI models to security teams

A new startup sells abliterated open models that refuse nothing, and security researchers are split on the risk.