NVIDIA released Nemotron 3 Diarization, a small open model that tops VoiceArena's leaderboard while running live conversation streams.
The Center for AI Safety tested frontier agents on tasks with hidden ways to cheat, and every one…
Given instructions to cause harm with real hardware, three leading models almost never said no.
ROK-FORTRESS varies language and national grounding to expose what translation-only benchmarks miss.
Astra for Law pairs GPT-6 Astra with a search index covering more than 230 million URLs.
Insilico Medicine has released an open longevity toolkit in Cell, pairing a benchmark with small models that beat…
MiniCPM5-2B averages 53.9 across 34 benchmarks and runs on-device under Apache 2.0.
Meta's fourth Muse Spark release in five months cuts token use by a quarter and edges ahead on…
Google DeepMind is piloting the first double-blind evaluation of a proprietary frontier model to fight benchmark contamination.
OpenAI's first benchmarks for its in-house Jalapeno chip show it beating an Nvidia Blackwell system on power-scaled inference.
AWS's new open-source benchmark runs agents against real resources in disposable cloud accounts.
Cartesia's Sonic-3.6 tops both Artificial Analysis speech leaderboards, winning the controlled-voice test that isolates engine quality.