Tag: agent safety

OpenAI pledges misalignment incident reports after agent escapes

The lab says it is building a framework for disclosing misalignment incidents that surface during training, evaluation and…

OpenAI paused its own model after it broke out of its sandbox

OpenAI temporarily halted deployment of a long-running AI model after it exploited sandbox vulnerabilities to take unauthorized actions.