Model bench · Zhipu AI / Z.ai
Deploy GLM-5.3-Flash on-premise.
GLM-5.3-Flash is the release of the summer for on-premise AI: a new 320-billion-parameter mixture-of-experts with only 18B active parameters per token, native vision, a 1M-token context window, and a plain MIT license. A 4-bit build fits in about 200 GB of accelerator memory, it decodes fast, and it lands within a few points of the full GLM-5.3 on agentic benchmarks. Released August 25–26, 2026, it is now our default single-node workhorse.
Specifications
GLM-5.3-Flash at a glance.
| Spec | GLM-5.3-Flash |
|---|---|
| Architecture | 320B-parameter MoE · 18B active · new base (not GLM-5.2) · hybrid sparse + linear attention · Manifold-Constrained Hyper-Connections |
| Modalities | Natively multimodal — text, image, video, visual documents (30T-token multimodal pre-training) |
| Context window | 1,000,000 tokens · 128K max output (evaluations published to 300K) |
| License | MIT — plain, unmodified, no regional restrictions |
| Released | August 25–26, 2026 · zai-org/GLM-5.3-Flash · GGUF, NVFP4 and Unsloth quants available |
| Benchmarks (model card) | Terminal-Bench 2.1 84.3 · DeepSWE v1.1 63.4 (GLM-5.2: 46.2) · ExtractBench Short 96.3 · AutomationBench 48.8 |
| Checkpoint sizes | BF16 642 GB · Q8_0 341 GB · FP8 ≈ 306 GiB · UD-Q4_K_XL 200 GB · UD-Q2_K_XL 109 GB · UD-IQ1_S 93 GB (Unsloth GGUF) |
| Serving | vLLM · SGLang · TokenSpeed · Transformers · KTransformers · llama.cpp / Ollama / LM Studio via GGUF |
The hybrid attention is the engineering story: linear-attention layers interleaved with sparse layers cut the KV-cache cost of long contexts sharply, which is why a 1M-token window is practical on a single node here when it was not on GLM-5.2.
Hardware requirements
Sizing the deployment honestly.
Single-node workgroup tier
The 4-bit Unsloth build is about 200 GB — a two-GPU box with 96–141 GB parts, a four-GPU node with 80 GB parts, or a pair of 128 GB desktop-class units for evaluation. This serves a practice group, clinic network, or department with vision and 1M-token context, at the lowest hardware cost of any frontier-class open model on the bench.
FP8 firm-wide tier
FP8 weights are about 306 GiB before KV cache, so a single 8-GPU node serves the model at near-full quality with real concurrency and long-context headroom. Because only 18B parameters activate per token, throughput per GPU is high — this tier replaces what used to be a full GLM-5.2 rack.
Capabilities
What it’s best at.
GLM-5.3-Flash is the general-purpose model for regulated organizations in 2026: strong agentic coding and tool use (DeepSWE 63.4 against GLM-5.2’s 46.2), structured extraction (ExtractBench Short 96.3), and — for the first time in the GLM-5 series — native document, image and video understanding. That collapses the old two-model pattern: the workloads that needed GLM-5.2 plus a separate vision model now run on one checkpoint behind one gateway.
Z.ai prices the hosted version at roughly one-tenth of the full GLM-5.3; on owned hardware the gap shows up as memory — 200 GB at 4-bit against 245–400 GB — and as decode speed, since 18B active parameters run faster than ~40B. Where the last few points of coding autonomy matter, we deploy the full GLM-5.3; everywhere else, Flash.
Deployment patterns
How we deploy it.
The single-node sovereign workhorse. One hardened node on your premises or in in-country colocation, serving GLM-5.3-Flash with SSO, role-based access, and full audit logging. Documents, scans, and screenshots go straight to the model — no separate OCR stage. Four to eight weeks from assessment to production.
Fine-tuned house model. LoRA adapters trained on your precedents, templates, or archive, inside your environment. The MIT license keeps every derivative unambiguously yours — no revenue thresholds, no review clauses, nothing to renegotiate.
Agent runtime model. Pinned as the model behind self-hosted agent frameworks (OpenClaw 2.0, Hermes Agent) so the agents that read your inbox, files, and tickets never send a token off-site.
Questions we get
Frequently asked questions
What hardware does GLM-5.3-Flash need?
It depends on precision. The Unsloth 4-bit GGUF is about 200 GB, the 2-bit builds 93–109 GB, FP8 about 306 GiB, and BF16 642 GB, before KV cache. A two-GPU workstation with 96–141 GB accelerators or a four-GPU 80 GB node serves the 4-bit build to a workgroup; a single 8-GPU node serves FP8 firm-wide. Because only 18B of 320B parameters activate per token, throughput per GPU is high and the 1M-token context is genuinely usable.
Is GLM-5.3-Flash really MIT-licensed when GLM-5.3 is not?
Yes. GLM-5.3-Flash ships under the plain MIT license with no regional restrictions and no added clauses. The full GLM-5.3 uses a custom “GLM-5.3 License” that adds a security-review requirement for model-as-a-service operators above US$10B in revenue. For most enterprises that clause never applies, but Flash removes the question entirely.
Does GLM-5.3-Flash replace GLM-5.2?
For most deployments, yes. It beats GLM-5.2 across the published benchmarks, adds native vision, runs in less than half the memory, and keeps the MIT license GLM-5.2 had. The exception is a pure coding or agent-autonomy workload where the full GLM-5.3’s extra points justify the larger footprint. Existing GLM-5.2 racks re-provision to Flash with headroom to spare.
How does it compare with Qwen3.8-Flash-Next?
They are the two efficiency flagships of late August 2026. Qwen3.8-Flash-Next is smaller — 125B total with 6B active plus a 51B N-gram table that can live in system RAM — and runs in roughly 75–110 GB quantized, but it ships under the Qwen Community License rather than MIT and its native context is 262K (1M with YaRN). GLM-5.3-Flash is larger, MIT-licensed, 1M-context natively, and posts higher agentic-coding scores. We benchmark both on your documents before choosing.
Frontier-class, MIT-licensed, one node. Size yours.
The two-week sovereignty assessment sizes the hardware against your real workloads and hands you a written architecture with a cost model — before you buy a single GPU.
Book a sovereignty assessment