Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks
Z.ai simply launched GLM-5.3. GLM-5.3 runs on the identical 743B base mannequin as GLM-5.2. Every reported achieve comes from scaled post-training: extra job environments, extra atmosphere sorts, longer coaching. The outcomes land in two locations. Coding jumps most on the longest-horizon benchmarks, with Terminal-Bench 3.0 shifting from 4.6 to twenty-eight.3. Cybersecurity moved additional than Z.ai says it anticipated, with CyberGym reaching 84.5%. Weights should not public but.
Is It Deployable?
Partially, GLM-5.3 is dwell by way of the Z.ai API, the GLM Coding Plan, and ZCode. Weights should not out. Z.ai says it would publish them roughly two weeks after launch, as soon as security analysis and hardening end.
- Which firms can transfer now: Startups and mid-market engineering orgs can undertake it at present through the Coding Plan or API. Enterprises with data-residency or vendor-review guidelines ought to look forward to weights. Security distributors and MSSPs get the most sign, and the most coverage publicity.
- Industries: Developer tooling, cloud infrastructure, software safety, fintech and e-commerce engineering, and distributors transport kernels, browser engines, or community stacks.
- Applications: Repository-scale refactors, long-horizon CLI brokers, CI failure triage, white-box vulnerability discovery, crash triage, and safe code evaluate.
Coding Results
Terminal-Bench 3.0 strikes from 4.6 to twenty-eight.3 in opposition to GLM-5.2. DeepSWE v1.1 strikes from 46.2 to 66.9. Agents’ Last Exam (CLI) strikes from 23.8 to twenty-eight.5. On GDPval-AA v2, which spans 44 occupations, GLM-5.3 scores 1,769.
On Z.ai Code Bench, an inside analysis, the firm reviews a 50% enchancment over GLM-5.2. It reviews 31.4% at roughly 50,000 output tokens per job. Claude Opus 4.8 scores 29.5% at 120,000 tokens. Claude Fable 5 nonetheless leads at 39.5% at most effort. Z.ai argues a non-public benchmark reduces contamination danger.
On public suites, GLM-5.3 trails GPT-5.6 Sol and Fable 5 on a number of more durable coding evaluations. All figures are vendor-reported, with harness, context size, and sampling settings documented in the announcement.
The Cybersecurity Result
Z.ai flags this one as unplanned. It added vulnerability-discovery knowledge anticipating higher single-bug reasoning. Instead, functionality stored compounding as coaching scaled. The mannequin started forming coherent plans throughout full exploitation chains.
CyberGym, which assessments discovery and validation from white-box supply, strikes from 77.2% to 84.5%. That edges previous Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. ExploitBench, which requires root-cause reasoning and a working exploit, strikes from 24.4% to 54.4%. Mythos 5 sits at 78.0%. On ExploitGym, GLM-5.3 completes 105 duties in two hours and 130 in six. GLM-5.2 completes 29 and 39. Mythos 5 completes 181 and 247.
The sample is constant. The deeper into the exploitation chain a benchmark sits, the bigger the achieve over GLM-5.2. The hole to closed frontier fashions additionally widens.
Interactive Explainer
Key Takeaways
- GLM-5.3 reuses the GLM-5.2 base mannequin; all beneficial properties come from post-training scaling.
- Terminal-Bench 3.0 strikes from 4.6 to twenty-eight.3; DeepSWE v1.1 from 46.2 to 66.9.
- CyberGym hits 84.5%, forward of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%).
- ExploitBench greater than doubles to 54.4%, however trails Mythos 5 at 78.0%.
- Weights ship in about two weeks, after security analysis and hardening.
Check out the Z.ai GLM-5.3 technical blog, Zai_org announcement, Z.ai Security Disclosure Ledger and zai-org/GLM-5 on GitHub. Also, be at liberty to observe us on Twitter and don’t neglect to hitch our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to companion with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and so forth.? Connect with us
The put up Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks appeared first on MarkTechPost.
