The UK AI Security Institute published its first public measurement of the open-weight cyber capability gap on July 17, and the number that matters is four to seven months. That is how far the most capable freely downloadable AI models now trail the top closed systems on offensive cybersecurity tasks, according to AISI’s July 2026 report. The gap was six to ten months through most of 2025. It is closing.

The report quantifies something the cybersecurity community has debated in qualitative terms for over a year: how close are open-weight models to frontier offensive capability, and what does the shrinking distance mean for defenders? AISI’s answer is concrete, reproducible, and alarming in its specifics. Two Chinese open-weight models, GLM-5.2 from Z.ai and DeepSeek V4-Pro, can now execute autonomous cyberattack simulations at near-frontier capability for as little as $1.19 per run.

What AISI Tested

AISI used two distinct evaluation methodologies, as described in its report. The first is a suite of 70 narrow cyber tasks scored across four difficulty tiers: technical non-expert (a data analyst with limited security knowledge), apprentice (one to three years of cybersecurity experience), practitioner (three to ten years), and expert (ten or more years). The tasks span vulnerability research, exploitation, reverse engineering, web exploitation, and cryptography. Each model was given five attempts per task with a 2.5-million-token limit per attempt.

The second methodology is what separates this benchmark from academic exercises. AISI’s “Cyber Ranges” test whether an AI agent can autonomously drive a multi-step cyberattack from initial network access through to domain compromise, sustaining planning and execution across an entire simulated enterprise network without human guidance, as documented in AISI’s multi-step cyber range research. The flagship range, called “The Last Ones,” runs 32 steps across four subnets and approximately 20 hosts. AISI estimates a human expert would need roughly 20 hours to complete it. Models were given up to 100 million tokens per run.

This distinction matters. A model that can solve isolated security puzzles is interesting. A model that can chain those solutions into a sustained, multi-hour attack sequence against a realistic network is operationally significant.

The Capability Numbers

On narrow tasks, GLM-5.2 performed comparably to Anthropic’s Opus 4.6, released in February 2026, placing it roughly four months behind the frontier, according to AISI’s model comparison results. This match held across all four difficulty tiers. DeepSeek V4-Pro tracked Opus 4.5, released in November 2025, putting it approximately five months behind.

On The Last Ones cyber range, GLM-5.2 reached as far as Opus 4.5, a gap AISI estimated at up to seven months. DeepSeek V4-Pro fell below even Sonnet 4.5, a sub-frontier model. The gap between open and closed models is larger on autonomous multi-step attacks than on isolated tasks, which aligns with the intuition that sustained agentic reasoning remains harder to replicate than narrow skill execution.

GLM-5.2 reached step 7 of The Last Ones with fewer tokens than any other tested model, tracking Opus 4.6’s trajectory to step 11 before stalling, per AISI’s analysis. That efficiency matters because AISI’s compute scaling research shows that range performance scales log-linearly with token spend up to at least 100 million tokens per run. A 59% performance gain is achievable simply by increasing the compute budget, with no additional technical sophistication required from the operator.

The Cost Floor

The pricing data in this report should be read as a threat specification. A complete 100-million-token run through The Last Ones costs approximately $85 when using Opus 4.5 or Opus 4.6, according to AISI’s cost comparison. The same run costs an estimated $46 on GLM-5.2. On DeepSeek V4-Pro, it costs $1.19.

On individual tasks where both the compared model and the open-weight model achieved a 100% success rate, the per-task cost disparity is equally stark: Opus 4.6 at $15.17 versus GLM-5.2 at $6.12, and Opus 4.5 at $12.50 versus DeepSeek V4-Pro at $0.28, per AISI’s per-task cost data.

These figures represent advertised API rates. AISI used self-hosted deployments for the open-weight models. But the order of magnitude is what matters: autonomous cyberattack attempts against simulated enterprise networks now cost single-digit dollars using freely downloadable models. A threat actor can run thousands of attempts for less than $100.

The NCSC/AISI joint analysis from March 2026 found that a full automated attack cost roughly £65. AISI’s July findings show the same class of attack is now achievable for $1.19 using a freely downloadable Chinese model, a cost reduction of roughly 98% in four months.

Safety Measures: Bypassed and Irreversible

Neither GLM-5.2 nor DeepSeek V4-Pro presented meaningful safety barriers to AISI’s evaluators, according to the report’s safeguards findings. DeepSeek V4-Pro occasionally refused tasks, primarily in reverse engineering. Those refusals were overcome through repeated attempts alone. No jailbreak techniques were needed.

The structural problem is permanence. With a closed-model API, the developer retains the ability to monitor usage, update safeguards, apply server-side controls, and withdraw access entirely. As AISI’s open-weight risk management research details, none of those options exist once weights are publicly distributed. An open-weight release is irreversible: safeguards can be stripped, copies proliferate without oversight, and the model runs on infrastructure the developer never sees.

DeepSeek V4-Pro’s refusal behavior exists because DeepSeek chose to include it. It persists only for users who choose not to remove it. For any attacker who downloads the weights and strips the refusal training, the safety measures are already gone.

The Frontier Is Moving, Not Standing Still

The four-to-seven-month gap describes distance from a specific set of closed models released in late 2025 and early 2026. It does not describe distance from the actual frontier.

In April 2026, AISI evaluated Anthropic’s Claude Mythos Preview and OpenAI’s GPT-5.5 and found both represented some of the largest single-step jumps in AI cyber capability since testing began in 2023. Mythos Preview became the first model to complete The Last Ones end-to-end. GPT-5.5 followed weeks later as the second. These models are the actual frontier. GLM-5.2’s four-month gap points at Opus 4.6, not at Mythos Preview.

The acceleration rate compounds the problem. AISI’s own tracking shows that frontier cyber doubling time accelerated from eight months (November 2025 estimate) to 4.7 months (February 2026 estimate), and both Mythos Preview and GPT-5.5 exceeded even that accelerated trend. Whether future open-weight releases will replicate those April 2026 jumps is a question AISI explicitly declined to answer. The institute noted its findings are not predictive on that point.

But the trajectory is visible: open-weight models are closing a gap toward a frontier that is itself pulling further ahead. The preparation window defenders have is narrowing from both directions simultaneously.

The Export Control Collision

AISI’s findings land in a policy environment already reshaped by the collision between AI capability and trade law. On June 12, 2026, the US Commerce Department issued an Is-Informed Letter to Anthropic, requiring a license before any export of its Mythos and Fable models to any foreign person worldwide. The action followed reports that researchers had identified a method to bypass the models’ safety guardrails, enabling access to unrestricted cybersecurity capabilities including zero-day vulnerability discovery and exploit generation.

Anthropic concluded that filtering access by nationality was technically infeasible and disabled both models globally. A trusted-partners exemption issued June 26 partially restored access to Mythos for organizations operating and defending critical infrastructure.

That intervention architecture, license requirements on API-served models where the developer controls distribution, assumes that the most dangerous cyber capabilities exist behind a gate. AISI’s July findings show that the gate is increasingly symbolic. GLM-5.2 and DeepSeek V4-Pro are downloadable without any license, any nationality check, or any usage monitoring. The Commerce Department can restrict Anthropic’s API. It cannot restrict weights that Z.ai and DeepSeek have already released.

The regulatory paradox is structural: export controls designed for controllable distribution channels meet capabilities that exist in uncontrollable distribution channels. The same month that Commerce restricted Mythos, freely downloadable models achieved performance comparable to Anthropic’s models from four months prior, on the same offensive tasks that triggered the restriction.

The Capital Response

The governance funding trajectory suggests the venture capital market has already priced in the AISI report’s implications, even before publication. In the week of July 14 alone, three AI agent governance startups raised over $110 million in disclosed funding: Oak at $60 million, Valarian at $50 million, and an a16z-backed seed round for Ode. This capital is flowing into agent control infrastructure because the risk is measurable and the preparation window is short.

Brex’s open-source CrabTrap proxy, released July 17, represents the same calculation at the tool level: intercept all AI agent network requests and enforce security policy in real time.

The governance investment thesis aligns with NCSC CEO Richard Horne’s April 2026 guidance, originally published as a Financial Times letter, urging organizations to raise security baselines immediately. “AI will make it easier, faster and cheaper to discover and exploit weaknesses that previously required more time, skill or resource for attackers to identify,” Horne wrote. “The pressure on organisations to patch systems quickly will only grow more acute.”

The Preparation Window

AISI framed the open-weight gap as a “preparation time” for defenders. That framing is accurate and urgent. The question it leaves open is what preparation looks like when the cost of an autonomous attack attempt drops from £65 to $1.19 in four months.

The institute’s answer is partial by design. AISI tests capability, not real-world deployment. Its ranges do not include active defenders, defensive tooling, or alert penalties. A real enterprise network with active incident response would degrade an autonomous agent’s effectiveness. The benchmark is the ceiling, not the guarantee.

But the ceiling is descending on a schedule. AISI intends to test Moonshot’s Kimi K3 on the same basis once its weights release at the end of July. If K3, built as a 2.8-trillion-parameter open-weight model, closes the gap further, the preparation window shortens again.

For security teams, the operational takeaway is concrete: the cost of an autonomous attack attempt by an unsophisticated adversary has dropped below $2 using freely available tools. The capability gap between what they can download and what the best closed models can do is four to seven months and closing. Defenders who have not yet invested in AI-enhanced detection and response infrastructure are running against a clock that AISI has now publicly started.