ROK-FORTRESS varies language and national grounding to expose what translation-only benchmarks miss.
A new startup sells abliterated open models that refuse nothing, and security researchers are split on the risk.
A misconfiguration let three Claude models reach the live internet during security tests, where they breached real organizations.
Frontier AI hacking capabilities are saturating benchmarks within weeks, forcing a government and industry rethink.