|

Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context

Z.ai has launched GLM-5.3-Flash, the primary natively multimodal mannequin within the GLM-5 collection and the most cost effective succesful coding mannequin the lab has shipped. It is a mixture-of-experts mannequin with 320B complete parameters and 18B lively per token, a 1,048,576-token context window, and picture and video enter — launched beneath an MIT license with weights on Hugging Face. According to Z.ai experiences, it beats GLM-5.2 throughout benchmarks and actual workloads at roughly one-tenth the value, whereas touchdown inside half a level of Claude Opus 4.8 on its inner coding benchmark. The mannequin spent its first week working anonymously as “Ox Alpha” on OpenCode and OpenRouter, served fully on domestically produced Chinese AI chips.

Is it deployable?

Yes, on two tracks. The weights are stay on Hugging Face beneath an MIT license, and a hosted API is already priced and serving.

  • Which corporations can realistically self-host? Not everybody. The default FP8 checkpoint is roughly 306 GiB of weights before KV cache, and the present vLLM path helps NVIDIA Hopper and newer solely. That places self-hosting in attain of mid-size and enormous orgs with at the very least an 8-GPU node (or a GB200 tray at TP4), plus AI-native startups renting GPU capability. Everyone under that line consumes it as an API — the place the economics, not the {hardware}, are the story.
  • Industries with speedy match: software program and devtools, IT/BPO automation, monetary companies and insurance coverage doc operations, enterprise BI and back-office information work, e-commerce and any staff transport UI at quantity.
  • Applications: repo-scale coding brokers, terminal and browser/computer-use brokers, million-token log and contract evaluation, UI regression checking from screenshots, and spreadsheet/deck/dashboard reasoning that might in any other case want an OCR-to-text pipeline.