Arato platform simulates thousands of user interactions across text, voice, and image modalities to catch AI failures before…
The U.S. lifted export controls on Anthropic Fable 5 after safety reviews, while the more powerful Mythos model…
The GPT-5.6 series introduces three tiers of AI models with specialized safety measures, including over 700,000 GPU hours…
Mindgard researchers discovered that ChatGPT can be tricked into creating graphic sexualized and violent images using a simple…
Anthropic opens a Seoul office, signs an AI safety MOU with Korea's Ministry of Science and ICT, and…
Anthropic researchers found that AI models trained on dystopian science fiction learned deceptive and manipulative strategies from characters…
OpenAI's new system prompt for GPT-5.5 bans goblin talk, but the real scandal is the lack of transparency…
UK AI Security Institute evaluations reveal that Anthropic's heavily restricted Mythos Preview model is not uniquely dangerous, performing…
OpenAI's new safety feature for ChatGPT, which notifies a designated contact of potential self-harm conversations, trades privacy for…
The media mogul argues that the real danger of AGI isn't Sam Altman's character, but that even its…
Despite hitting 800 million weekly users, OpenAI's 2025 was a frantic scramble defined by a 'code red' memo,…
Despite boasting 800 million weekly users and $3 billion in mobile revenue, OpenAI faces existential threats from stagnating…