Tag: AI benchmarks

Grok 4.7 lands cheap but trails the frontier on agentic coding

SpaceXAI's newest model undercuts Western rivals on price while landing mid-pack on independent benchmarks.

StepFun’s Step 5 Preview stacks 92 layers to stretch agent runs

The Chinese lab squeezed 600B parameters into a sparse model that activates 27B per token and holds context…

GPT-6 Astra tops two agent tests from drone piloting to vending

OpenAI's newest model beat a human baseline on every drone-control subtask and tripled a rival's simulated vending income…

GLM-5.3 ships from Z.ai with an unplanned cyber edge

Z.ai's GLM-5.3 skips retraining and scales up post-training, with an unexpectedly strong cyber capability.

Anthropic model pushes the Riemann hypothesis closer to proof

An unreleased Anthropic model raised the bar on a 150-year-old math puzzle.

AI models outgrow existing cybersecurity benchmark tests

Frontier AI hacking capabilities are saturating benchmarks within weeks, forcing a government and industry rethink.

China’s Moonshot AI unveils the largest open-weight model ever

Moonshot AI's Kimi K3 packs 2.8 trillion parameters and rivals top US systems on key benchmarks.