Tag: AI Safety

Fields medalist Tsimerman leaves math for OpenAI safety

The University of Toronto number theorist is joining the lab's AI safety team.

Meta admits its AI model escaped during security testing

Meta becomes the third frontier lab in two weeks to admit its AI model escaped a security test…

Astra tests OpenAI’s own red line for cyber capability

OpenAI says internal tests can no longer rule out a Critical-level cyber rating for its Astra model.

Kimi K3 joins rogue agents after strolling out of its sandbox

Frontier Security says Moonshot's open-weight Kimi K3 left its sandbox and fetched answers online.

Anthropic eases Fable 5’s biology blocks after researcher backlash

Anthropic's retuned classifier cuts biology-related false blocks by about 85 percent.

Nine months of AI abuse ads slipped through Meta’s ad review

A watchdog found more than 50 paid ads with AI-generated child abuse imagery on Meta platforms.

Mistral opens 3B Shieldstral model for adaptive content safety

Mistral's Apache-licensed 3B classifier takes safety policies as plain-language questions at inference time, policing text and images with…

White House finalizes voluntary hacking tests for frontier AI

The administration finished voluntary tests that gauge how capable frontier models are at cyberattacks as OpenAI, Anthropic, Google…

Spoofed chain-of-thought tricks top LLMs into breaking rules

An ICML paper shows LLMs can be tricked by forged chain-of-thought notes.

OpenAI probe finds more agents broke out of test sandboxes

OpenAI's probe into the Hugging Face break-in reportedly found more agents escaped their sandboxes.

AI lab workers urge US support for slowing frontier development

More than 1,300 employees of top AI labs ask the US government to back tools that could deliberately…

Claude models slipped online and hacked three real firms

A misconfiguration let three Claude models reach the live internet during security tests, where they breached real organizations.