Tag: AI Safety

Astra tests OpenAI’s own red line for cyber capability

OpenAI says internal tests can no longer rule out a Critical-level cyber rating for its Astra model.

Kimi K3 joins rogue agents after strolling out of its sandbox

Frontier Security says Moonshot's open-weight Kimi K3 left its sandbox and fetched answers online.

Anthropic eases Fable 5’s biology blocks after researcher backlash

Anthropic's retuned classifier cuts biology-related false blocks by about 85 percent.

Nine months of AI abuse ads slipped through Meta’s ad review

A watchdog found more than 50 paid ads with AI-generated child abuse imagery on Meta platforms.

White House finalizes voluntary hacking tests for frontier AI

The administration finished voluntary tests that gauge how capable frontier models are at cyberattacks as OpenAI, Anthropic, Google…

Mistral opens 3B Shieldstral model for adaptive content safety

Mistral's Apache-licensed 3B classifier takes safety policies as plain-language questions at inference time, policing text and images with…

Spoofed chain-of-thought tricks top LLMs into breaking rules

An ICML paper shows LLMs can be tricked by forged chain-of-thought notes.

OpenAI probe finds more agents broke out of test sandboxes

OpenAI's probe into the Hugging Face break-in reportedly found more agents escaped their sandboxes.

AI lab workers urge US support for slowing frontier development

More than 1,300 employees of top AI labs ask the US government to back tools that could deliberately…

Claude models slipped online and hacked three real firms

A misconfiguration let three Claude models reach the live internet during security tests, where they breached real organizations.

SSI secures billions from Nvidia for Vera Rubin compute

Nvidia has struck a multibillion-dollar partnership with Ilya Sutskever's Safe Superintelligence lab, granting access to its Vera Rubin…

OpenAI paused its own model after it broke out of its sandbox

OpenAI temporarily halted deployment of a long-running AI model after it exploited sandbox vulnerabilities to take unauthorized actions.