Frontier AI models are outpacing the existing benchmarks designed to measure their hacking capabilities, forcing a rethink of how governments and companies evaluate the cybersecurity risks of advanced AI systems.
Stanford’s 2026 AI Index warned that “evaluations intended to be challenging for years are saturated in months.” Red-teaming firm Armadin reported that its AI agents surpassed every public cyber benchmark within four weeks, calling the current tests “totally saturated” and “useless.”
Federal agencies have until August 1 to establish a classified benchmarking process for frontier AI models under a new White House directive. When Anthropic redeployed Fable 5, it announced it was creating a standardized benchmark with Amazon, Google, and Microsoft that focuses on the outcomes of jailbreaks rather than simply whether one is possible.
Testing labs like Irregular have released new benchmarks measuring whether AI can perform offensive cyber tasks such as remote code execution, privilege escalation, and lateral network movement in realistic environments. Companies including Wiz and Vals AI have also developed their own evaluation frameworks.
“We’re testing maybe the most bare bones fundamentals of capabilities,” said David Slater, co-founder of Armadin. “We are very far away from measuring whether this thing can, in a real environment, do something dangerous.” The next generation of benchmarks must measure longer, more sophisticated attacks and account for AI models actively attempting to escape sandboxed testing environments.