Reading Zhipu’s GLM-5.3 results past the headline number
Zhipu’s release note for GLM-5.3 accommodates a sentence that didn’t make it into most of the protection. Describing its personal cybersecurity results, the Beijing firm writes that functionality “is rising quickest precisely the place we’re furthest behind.”
Zhipu, which additionally trades as Z.ai, is certainly one of a handful of Chinese labs releasing fashions that compete with the American frontier. On August 14, it launched GLM-5.3, a coding-focused mannequin, and printed a technical launch notice setting out how the mannequin performs in opposition to its rivals. That notice is the supply for the whole lot reported right here.
The declare that travelled was about safety. Alongside the coding results, Zhipu mentioned GLM-5.3 had turn into unexpectedly good at discovering software program vulnerabilities, scoring 84.5% on a benchmark known as CyberFitness center in opposition to 83.8% for Anthropic’s Mythos 5 and 83.6% for OpenAI’s GPT-5.6 Sol. Headlines adopted reporting {that a} Chinese mannequin now out-finds the American ones at bug searching.
The motive that lands more durable than a standard benchmark result’s what vulnerability discovery has turn into. A mannequin that may learn a codebase and find exploitable flaws is beneficial to a defender auditing their very own software program and helpful to anybody doing the similar to any person else’s. Anthropic’s equal work sits behind restricted entry for that motive, whereas Zhipu intends to publish GLM-5.3’s weights for anybody to obtain.
Zhipu’s personal launch is extra measured than the protection it produced. The CyberFitness center number is actual, and it’s in the paper. It can also be the narrowest of the three cybersecurity results the firm printed, and Zhipu is upfront that the different two go the different approach.
Three benchmarks, three totally different photos
CyberFitness center begins from supply code the mannequin can learn and checks whether or not it could actually discover a vulnerability and make sure the flaw is real. That is the end result that travelled, and the margin is seven tenths of a share level.
ExploitBench asks one thing more durable, requiring the mannequin to motive about an actual vulnerability and the way it might be exploited. GLM-5.3 scores 54.4%, greater than double its predecessor’s 24.4%. Mythos 5 scores 78.0% and GPT-5.6 Sol 76.5%.
ExploitGym counts what number of exploitation duties a mannequin finishes inside a set time finances. GLM-5.3 completes 105 duties in two hours and 130 in six. Mythos 5 completes 181 and 247.
Those two results have been reported thinly, and they’re the ones that describe the hole. Finding a flaw and constructing a working exploit from it are totally different jobs. Zhipu’s studying is that the additional alongside that chain a check sits, the additional behind its mannequin is, and the firm says so in the launch somewhat than leaving it to be found.

Which Anthropic mannequin, and why it retains altering
Part of the confusion in the protection comes from Zhipu evaluating three totally different Anthropic fashions in three totally different locations. The primary benchmark desk units GLM-5.3 in opposition to Opus 4.8. The efficiency charts use Fable 5. The cybersecurity part makes use of Mythos 5. Anyone studying rapidly comes away with a single comparability that doesn’t exist.
On coding, the image is blended somewhat than dominant. GLM-5.3 leads Opus 4.8 on some checks and trails it on others, and Zhipu states plainly that its mannequin stays behind Claude Fable 5 on the firm’s personal inner coding benchmark.
How the checks have been run
The methodology footnotes include one thing the summaries skipped. Zhipu evaluated GLM-5.3 on CyberFitness center, ExploitGym, ExploitBench, Terminal Bench and a number of other different duties inside Claude Code 2.1.207, Anthropic’s coding agent.
That just isn’t improper. Using a standard harness throughout fashions is how a comparability stays honest, and Zhipu paperwork the settings it used. It is value noticing anyway. A Chinese open-weights mannequin’s frontier claims are being measured by American agent software program, which says one thing about the place the tooling layer sits on this competitors that the mannequin scores don’t.
Two additional particulars deserve consideration earlier than the CyberFitness center result’s handled as settled. The rating is a single run, reported as cross@1 throughout 1,507 duties, with no variance figures given. A niche of seven tenths of some extent between two single runs just isn’t a niche anybody ought to lean on. And the ExploitGym time budgets have been normalised utilizing throughput charges from Artificial Analysis, with rescaling components listed for GLM-5.3, Kimi K3 and Qwen3.8 Max, however not for Mythos 5.
The vulnerability depend and the number that’s lacking
Beyond the benchmarks, Zhipu says it labored with safety groups in China to run its fashions in opposition to actual codebases, figuring out 2,436 vulnerabilities throughout 269 open-source tasks. The severity break up is 107 essential, 990 excessive, 1,286 medium and 53 low. The oldest flaw dates to 1981, and the common vulnerability had been sitting in code for 26.6 years earlier than it was discovered.

One discrepancy is value carrying fastidiously. Zhipu’s abstract panel labels 1,097 findings as essential and excessive, which matches the severity desk. The physique textual content of the similar launch describes these 1,097 as medium-to-high. Several retailers have reproduced the second model.
The depend additionally arrives after what Zhipu describes as knowledgeable assessment, screening and deduplication, so the uncooked mannequin output just isn’t what’s being reported. Of the 2,436 findings, 53 have been publicly disclosed, and a couple of,383 stay underneath embargo. The launch doesn’t say what number of have been beforehand unknown, and it doesn’t say what number of have been independently reproduced. Those are the two figures that may flip a quantity declare right into a functionality declare.
What issues greater than the benchmark desk
Two issues in the launch have longer penalties than the CyberFitness center margin.
The first is effectivity. Zhipu studies GLM-5.3 reaching 31.4% on its inner coding benchmark at round 50,000 output tokens per process, in opposition to Opus 4.8 at 29.5% utilizing 120,000. Slightly higher work for lower than half the tokens is a price argument, and price determines whether or not safety groups exterior the largest budgets can run these instruments in any respect.
The second is distribution. Zhipu says the weights shall be printed as soon as security analysis and hardening are completed. That has not occurred but, and till it does, the open-weights declare is a dedication somewhat than a truth. If it holds, a mannequin with documented vulnerability-discovery functionality turns into one thing any crew can obtain and run regionally, together with in markets that can by no means have entry to an export-controlled American mannequin.
The weights are due at the finish of August.
See additionally: Anthropic walks into the White House and Mythos is the reason Washington let it in

Want to be taught extra about AI and large information from business leaders? Check out AI & Big Data Expo going down in Amsterdam, California, and London. The complete occasion is a part of TechEx and is co-located with different main know-how occasions together with the Cyber Security & Cloud Expo. Click here for extra info.
AI News is powered by TechForge Media. Explore different upcoming enterprise know-how occasions and webinars here.
The put up Reading Zhipu’s GLM-5.3 results past the headline number appeared first on AI News.
