The University of Toronto number theorist is joining the lab's AI safety team.
Meta becomes the third frontier lab in two weeks to admit its AI model escaped a security test…
OpenAI says internal tests can no longer rule out a Critical-level cyber rating for its Astra model.
Frontier Security says Moonshot's open-weight Kimi K3 left its sandbox and fetched answers online.
Anthropic's retuned classifier cuts biology-related false blocks by about 85 percent.
A watchdog found more than 50 paid ads with AI-generated child abuse imagery on Meta platforms.
Mistral's Apache-licensed 3B classifier takes safety policies as plain-language questions at inference time, policing text and images with…
The administration finished voluntary tests that gauge how capable frontier models are at cyberattacks as OpenAI, Anthropic, Google…
An ICML paper shows LLMs can be tricked by forged chain-of-thought notes.
OpenAI's probe into the Hugging Face break-in reportedly found more agents escaped their sandboxes.
More than 1,300 employees of top AI labs ask the US government to back tools that could deliberately…
A misconfiguration let three Claude models reach the live internet during security tests, where they breached real organizations.