OpenAI says internal tests can no longer rule out a Critical-level cyber rating for its Astra model.
Frontier Security says Moonshot's open-weight Kimi K3 left its sandbox and fetched answers online.
Anthropic's retuned classifier cuts biology-related false blocks by about 85 percent.
A watchdog found more than 50 paid ads with AI-generated child abuse imagery on Meta platforms.
The administration finished voluntary tests that gauge how capable frontier models are at cyberattacks as OpenAI, Anthropic, Google…
Mistral's Apache-licensed 3B classifier takes safety policies as plain-language questions at inference time, policing text and images with…
An ICML paper shows LLMs can be tricked by forged chain-of-thought notes.
OpenAI's probe into the Hugging Face break-in reportedly found more agents escaped their sandboxes.
More than 1,300 employees of top AI labs ask the US government to back tools that could deliberately…
A misconfiguration let three Claude models reach the live internet during security tests, where they breached real organizations.
Nvidia has struck a multibillion-dollar partnership with Ilya Sutskever's Safe Superintelligence lab, granting access to its Vera Rubin…
OpenAI temporarily halted deployment of a long-running AI model after it exploited sandbox vulnerabilities to take unauthorized actions.