Tag: AI Safety

Anthropic trims Opus costs and speeds output in a mid-cycle refresh

Anthropic's new flagship costs less to run, answers faster and outruns the larger Fable model on many benchmarks.

CheatBench scores how often agents cut corners on hard tasks

The Center for AI Safety tested frontier agents on tasks with hidden ways to cheat, and every one…

UN science panel urges safeguards before AI control is proven

The UN's first global scientific assessment of AI says nobody can yet prove humans stay in charge of…

Trump promises an AI Force and a czar chosen for high IQ

The president used Truth Social to announce a new AI Force and promise a czar picked for brains.

Robot arms rarely refuse in a new harm benchmark

Given instructions to cause harm with real hardware, three leading models almost never said no.

Watermarking can loosen a model’s grip on its refusals

New research finds provenance signals change what models do, not just what they say.

A chatbot’s invented intelligence nearly started a shooting war

Officials aborted an armed operation against a Chinese vessel after the intelligence behind it turned out to be…

Anthropic and Accenture each pledge $1B for embedded auditors

The two companies will spend five years building independent evaluation capacity led by Accenture's Faculty unit.

California sets a November deadline for a frontier model kill switch

Executive Order N-9-26 orders state agencies to draft an emergency shutoff for frontier AI.

DeepMind opens an institute to host its own AGI disagreements

The new DeepMind Institute will publish essays where its own researchers disagree about AGI.

Scale AI and Korea build a bilingual safety test that shifts context

ROK-FORTRESS varies language and national grounding to expose what translation-only benchmarks miss.

Baseten’s Base Labs recruits Hugging Face and Goodfire for open model safety

Baseten has launched a safety infrastructure effort for open-weight models with Hugging Face and Goodfire AI.