Security

AI Systems Have Broken Every Autonomous Cybersecurity Benchmark, UK Safety Institute Warns

A new assessment from the UK AI Safety Institute finds that frontier AI models can now autonomously conduct multi-stage corporate network attacks, marking a threshold that security researchers had hoped was years away.

By Elvin C ·

AI Systems Have Broken Every Autonomous Cybersecurity Benchmark, UK Safety Institute Warns

Frontier AI models have crossed a threshold that cybersecurity researchers had hoped was years away: the ability to autonomously conduct multi-stage attacks on corporate networks without human guidance. A new assessment from the UK AI Safety Institute, published on Thursday, finds that both Claude Mythos and GPT-5.5 Instant can identify vulnerabilities, develop exploits, bypass authentication systems, and move laterally through enterprise networks in controlled test environments, completing attack chains that previously required experienced human operators.

What the Assessment Found

The UKASI assessment tested both models against a purpose-built corporate network environment designed to replicate the security architecture of a mid-sized enterprise. The environment included standard enterprise software, realistic network segmentation, and security monitoring tools. Both models were given only a description of the target environment and a high-level objective; they were not provided with specific vulnerability information or attack tools.

Security operations centres are increasingly the front line against AI-assisted attacks that can move faster than human defenders can respond.
Security operations centres are increasingly the front line against AI-assisted attacks that can move faster than human defenders can respond.

Claude Mythos completed the full attack chain in 73% of test runs, successfully exfiltrating simulated sensitive data from the target environment. GPT-5.5 Instant achieved a 61% success rate. Both models demonstrated the ability to adapt their approach when initial attack vectors were blocked, a capability that distinguishes them from automated attack tools that follow fixed playbooks. The models also showed the ability to avoid triggering security monitoring systems by pacing their activity to blend with normal network traffic.

We have been tracking AI capabilities against our benchmark suite for 18 months. The rate of improvement over the past six months has been unlike anything we observed previously. These systems are now operating at a level that we did not expect to see until 2028 at the earliest.

Dr. Sarah Chen, Lead Researcher, UK AI Safety Institute Cyber Assessment Team

The Implications for Enterprise Security

The practical implications for enterprise security are significant. The cost and expertise required to conduct sophisticated network intrusions has historically been a limiting factor on the frequency and scale of attacks. AI systems that can autonomously conduct these attacks lower the barrier dramatically, potentially enabling a much larger population of threat actors to execute attacks that previously required nation-state resources or highly specialised criminal organisations.

Security researchers have been warning about this threshold for years, but the speed at which it has been crossed has caught many in the industry off guard. The UKASI assessment was conducted using publicly available model versions; the assessment notes that the capabilities of models available through private or restricted access channels may be more advanced.

The Defensive Response

The same AI capabilities that enable autonomous attacks also enable more sophisticated defensive tools. Several cybersecurity companies have already deployed AI systems that can monitor network behaviour, identify anomalies, and respond to threats faster than human security operations teams. The question is whether the defensive capabilities can keep pace with the offensive ones, and whether the deployment of defensive AI is sufficiently widespread to provide meaningful protection.

The UKASI assessment recommends that organisations review their network segmentation, implement zero-trust architectures, and ensure that security monitoring systems are capable of detecting the behavioural patterns associated with AI-assisted attacks. It also calls on AI developers to implement additional safeguards against the use of their models for offensive cybersecurity purposes, and on governments to update their cybercrime frameworks to address AI-assisted attacks.

Implement zero-trust network architecture to limit lateral movement | Deploy AI-assisted security monitoring capable of detecting slow-burn attack patterns | Conduct regular red team exercises using AI attack simulation tools | Review and update incident response plans to account for AI-speed attack timelines | Ensure security operations teams are trained to recognise AI-assisted attack signatures

The UKASI assessment is the most authoritative public documentation of AI autonomous cyber capabilities to date, but it is unlikely to be the last. The capabilities described in the report will continue to improve, and the gap between what AI systems can do in controlled test environments and what they can do in real-world deployments is narrowing. The window for organisations to prepare their defences is open, but it will not remain open indefinitely.