Tag: alignment

Claude-built researchers beat humans at fixing model flaws

Anthropic says automated alignment agents closed more of the safety gap than expert humans on seven of seven…

Fields medalist Tsimerman leaves math for OpenAI safety

The University of Toronto number theorist is joining the lab's AI safety team.