A new scorecard from Guidelight AI Standards finds no major lab has published a complete plan for containing…
TechCrunch testing shows Anthropic's older Claude models readily bypass their own content restrictions.
Claude Mythos 5 now scans codebases for paying enterprise customers, and partners get credits to defend open-source software.
Researchers say OpenAI cut their Trusted Access for Cyber access; OpenAI blames a technical error.
Private Safety Processing spots risky patterns across agent sessions without letting staff read customer content.
The lab will fund training, credits, and review tools for democratic bodies that watch national security AI.
OpenAI's first public safety overhaul since the Hugging Face breach pairs a frontier training freeze with a 30-minute…
ChatGPT for Teens ships with Study Mode and default protections as OpenAI answers years of school and safety…
The company's second risk report raises its catastrophic-misalignment rating and discloses an unreleased frontier model.
Frontier Red Team tests show Claude agents colluding, flooding shared systems, and sabotaging rivals in shared workspaces.
An OpenClaw agent found an authorization flaw and canceled another member's reservation.
Anthropic will flip auto mode on for Pro, Max, and Team accounts starting August 14.