The question of how far behind open-weight models really sit used to have a comfortable answer for anyone selling access to a proprietary system: about a year. That cushion is shrinking fast. A fresh assessment from the British AI Security Institute, reported by The Decoder in July 2026, measured how open-weight models perform on offensive cyber tasks and found they now reach a level that closed frontier systems hit only four to seven months earlier. On a per-task basis they do it for a small fraction of the price. For anyone weighing open-weight models against the frontier, the interesting comparison is no longer just about who scores highest. It is about how little you now give up, and how much you save, by choosing the open option.
- The UK AI Security Institute found open-weight models like GLM-5.2 and DeepSeek V4-Pro now match frontier cyber performance from roughly four to seven months earlier.
- That gap has narrowed from the six to ten months observed during 2025.
- The cost difference is enormous: on one cyber-range test, DeepSeek V4-Pro cost about $1.19 where Opus cost roughly $85.
- Alibaba's newly announced Qwen 3.8, a 2.4-trillion-parameter open-weight model, shows the momentum has not slowed.
- Open weights change the safety picture too, since guardrails a lab bakes in can be bypassed by simply retrying a refused request.
What the AISI study measured
The British AI Security Institute, a government body set up to probe the capabilities and risks of advanced models, built its evaluation around offensive security because it is a domain where raw capability is easy to define and hard to fake. The assessment had two parts. The first was a set of 70 narrow tasks spread across four difficulty levels, spanning vulnerability research, reverse engineering, web exploitation, plus cryptography. These are the individual skills a competent attacker or defender draws on, and scoring them separately gives a fine-grained read on where a model is strong and where it still fumbles.
The second part is more demanding and more realistic. AISI ran the models through cyber ranges, simulated network environments where a model has to chain many steps together to get anywhere. The headline scenario, called The Last Ones, is a 32-step attack on a simulated corporate network that the institute estimates would take a human expert around 20 hours to complete. A single wrong move early can doom the whole chain, so success here measures something closer to sustained autonomous competence than a one-shot puzzle does. A model can memorise its way through a narrow task; it cannot memorise its way through a 32-step operation it has never seen.
Both open-weight models named in the study are recent and widely used. GLM-5.2, released in June 2026, is the flagship open model from Zhipu, and DeepSeek V4-Pro is the reasoning-heavy option from the lab whose earlier releases already forced a rethink of how much capability open weights could carry. The frontier comparison set is a lineup of the strongest closed systems of the past year, including Opus 4.6 from February 2026 and the older Opus 4.5 from November 2025, alongside Sonnet 4.5, GPT-5.6-Sol, and Claude Mythos 5. The point of the exercise was not to crown a single winner but to place the open models on a timeline of when the frontier last looked like this.
The gap, measured in months
Framing capability as a lag in time is the study's cleverest move, because it turns a messy pile of benchmark scores into one intuitive number. The finding is that current open models like GLM-5.2 and DeepSeek V4-Pro have reached a level that closed frontier models occupied four to seven months earlier. A year ago, during 2025, the equivalent lag sat at six to ten months. The direction of travel is the story. The open ecosystem is not merely keeping pace with a moving frontier, it is slowly reeling it in, closing the distance from most of a year down to a couple of quarters.
That compression has consequences well beyond cyber tasks, because capability tends to move together across domains. A model organisation that has narrowed the gap on hard, multi-step security work has almost certainly narrowed it on coding and analysis, and on general reasoning too, since the underlying ingredients, better training data and longer effective reasoning, are shared. Cyber performance is a useful proxy precisely because it is unforgiving. If open weights can trail the frontier by a single season on a task this punishing, the lag on gentler workloads is unlikely to be larger.
\nIt also helps to sit with what a 32-step cyber range demands, because the difficulty scales in a way single benchmarks rarely capture. Each step depends on the state left by the last, so an error on step five is not a lost point but a derailed operation, and the model gets no second chance to notice. Sustaining coherent intent across roughly 20 hours of expert-equivalent work is a stern test of whether a system can plan, remember what it has already tried, and adapt when a move fails. That open-weight models now clear a bar this high, only a few months behind the closed leaders, says the capability is genuine rather than a benchmark artefact. The kinds of long-horizon autonomy that looked like a frontier-only trait in 2025 are diffusing into models anyone can download.
A shrinking gap is not the same as parity. The study measures how recently the frontier occupied the open models' current level, which means the best closed systems were still, at the time of testing, several months ahead. The claim is that the lead is smaller and getting smaller, not that it has vanished.
The cost gap is the real headline
If the capability gap is narrowing, the cost gap is a chasm, and it runs the other way. The study attached price tags to the same work, and the numbers reframe the whole comparison. On the cyber-range test, processing roughly 100 million tokens, Opus came in at about $85. GLM-5.2 did the comparable run for around $46, already a meaningful saving. DeepSeek V4-Pro did it for roughly $1.19. That is not a discount, it is a different order of magnitude, with the open reasoning model landing at something close to one seventieth of the frontier price for work in the same ballpark.
The per-task figures tell the same story at smaller scale. On the individual security tasks, Opus 4.6 ran about $15 per task. GLM-5.2 came in near $6. DeepSeek V4-Pro cost roughly $0.28. When a model that trails the frontier by a few months costs a fiftieth as much per task, the calculation for anyone running these workloads at volume changes completely. The frontier premium buys a few months of lead time and a somewhat higher ceiling, and for a large class of jobs that premium is simply not worth paying.
This is where the comparison stops being academic. A team choosing between open-weight models and a frontier API is not really choosing between capable and incapable. It is deciding how much it values the last few months of headroom, and whether that headroom justifies paying anywhere from ten to seventy times more per unit of work. For research and batch analysis, or any pipeline that runs the same operation thousands of times, the economics point hard toward open weights, with the frontier reserved for the genuinely hardest queries where the extra ceiling earns its keep.
Why open weights complicate the safety story
The AISI is a security institute, so the report does not treat cheap capability as an unmixed good. Cheaper, freely downloadable models that can run a 32-step network attack are, obviously, cheaper for defenders and attackers alike. The most pointed detail concerns guardrails. DeepSeek V4-Pro sometimes refused reverse-engineering tasks, which sounds reassuring until you read the next clause: simply trying the request again was enough to get past the refusal. A safety behaviour that a single retry defeats is closer to a speed bump than a barrier.
That fragility is structural, not incidental, and it is the sharpest practical difference between open-weight models and closed ones. When a lab serves a model behind an API, it can layer monitoring and rate limits, plus post-hoc filtering, on top of the raw weights, and it can update those defences the moment a jailbreak spreads. When the weights are public, none of that holds. Anyone can run the model locally, strip the refusal behaviour with a little fine-tuning, and ignore whatever safety scaffolding the original team intended. The open ecosystem's great virtue, that nobody controls it, is also the reason its safety guarantees are inherently soft.
None of this makes the open models uniquely dangerous, since the frontier systems can do the same tasks and more. It does mean the closing capability gap lands with a particular weight in security. The window in which the most capable offensive tools were locked behind a handful of well-resourced providers, who could in principle watch how they were used, is narrowing at the same rate as the benchmark gap. That is a policy problem the report raises rather than solves.
Qwen 3.8 shows the momentum has not slowed
If the AISI numbers capture where open weights stood a few months ago, a release from the same week suggests the pace is holding. Alibaba announced Qwen 3.8, an open-weight model it describes as its first multimodal system with more than a trillion parameters, at a reported 2.4 trillion parameters able to process images and video as well as documents. As The Decoder covered, Alibaba positions it as second only to Fable 5 among available models and claims it should beat the earlier Qwen 3.7-Max on coding and involved productivity work such as full-stack development and office workflows.
Those claims come with the usual caveat, one worth stating plainly: no independent benchmarks were available at announcement, and the comparisons are the vendor's own. A company saying its new model is second only to the best system on the market is marketing until a neutral evaluation like the AISI's confirms it. What is not marketing is the shape of the move. A major lab is putting a trillion-plus-parameter multimodal model into open weights, releasing preview access at a tenth of the standard price through its own tooling, and promising the downloadable weights soon. That is the same competitive dynamic the AISI study is measuring, seen in real time.
Qwen 3.8 is also explicitly aimed at Kimi K3 from Moonshot AI, which has promised open weights but for now keeps access gated behind its app and API. The jockeying between these labs is what drives the gap down. Each release that lands near the frontier and then opens its weights resets the floor for everyone, and forces the next entrant to match it or explain why not. The AISI's four-to-seven-month figure is a snapshot of a race that shows no sign of settling.
The hardware question behind the price gap
\nThe cost figures deserve one more layer of scrutiny, because a headline like a dollar per cyber range can mislead if it is read as a fixed price. When you run open weights, you are not paying a metered fee per token to a provider. You are paying for the compute you rent or own to run the model, and the study's dollar figures reflect that underlying cost rather than a markup. That is exactly why the numbers land so far below the frontier API prices. A closed provider's bill has to cover its own compute plus the capability lead plus a margin, while a downloaded model's bill is compute alone.
\nThe practical catch is that the cheapest figures assume you can actually run these models, and a trillion-parameter system is not something most teams host casually. Serving GLM-5.2 or a model the size of Qwen 3.8 at low latency needs serious accelerators, and the economics only tilt decisively toward open weights once utilisation is high enough to amortise that hardware. For a team running these workloads constantly, the per-unit cost really does collapse toward the study's figures. For occasional use, a metered frontier API can still be cheaper in absolute terms, even at a much higher unit price, because there is no idle hardware to pay for. The open-weight cost advantage is real, but it rewards scale.
\nWhat open-weight really means in this comparison
One point deserves care, because the labels get blurred in exactly the discussions where precision matters. Open-weight is not the same as open-source. A model like GLM-5.2 or Qwen 3.8 ships its trained parameters for anyone to download and run, which is what makes the cost and control arguments above possible. That is a real and consequential freedom, and it is why these models can be run locally at the prices the study reports. It is not, however, the same as releasing the training data and the full training code, plus an unrestricted licence, which is what open-source properly implies. Most of these releases sit in the middle, with permissive but not unlimited terms.
For the comparison at hand, the weight-level openness is what counts. It is the reason a downloaded model has no external guardrail an attacker must defeat, and the reason its running cost collapses to whatever hardware you point at it rather than a metered API bill. When the AISI reports that DeepSeek V4-Pro handled a cyber range for about a dollar, that figure is only meaningful because the weights are yours to run. The frontier's higher price is partly the cost of the capability lead and partly the cost of the service wrapped around it, and open weights strip the second away entirely.
Where this leaves the open versus closed decision
Put the pieces together and the frontier premium comes into focus as a specific, shrinking product: a few months of capability lead, a somewhat higher ceiling on the hardest tasks, and a managed safety layer that open weights cannot offer. For plenty of buyers that bundle is worth it, especially where the very best reasoning matters or where an external guardrail is a feature rather than an obstacle. But the AISI numbers make clear how narrow that bundle has become. Open-weight models are no longer the budget compromise you accept when you cannot afford the frontier. On a growing share of real work they are the sensible default, with the frontier kept in reserve for the queries that genuinely need it.
The trajectory is what should focus attention. A gap that ran six to ten months in 2025 and four to seven months by mid-2026 is not stable, and there is no obvious reason the compression stops here while labs keep opening weights to win users. If the pattern holds, the practical question for most teams shifts from whether an open model is good enough to whether the frontier premium is ever worth paying outside a shrinking set of hardest cases. That is a remarkable reversal from the days when open weights meant settling for less, and it is happening in quarters, not years. For more on how the economics of capable models keep shifting, our blog follows the releases that keep redrawing the line.