Model bench · Zhipu AI / Z.ai

Deploy GLM-5.3-Flash on-premise.

GLM-5.3-Flash is the release of the summer for on-premise AI: a new 320-billion-parameter mixture-of-experts with only 18B active parameters per token, native vision, a 1M-token context window, and a plain MIT license. A 4-bit build fits in about 200 GB of accelerator memory, it decodes fast, and it lands within a few points of the full GLM-5.3 on agentic benchmarks. Released August 25–26, 2026, it is now our default single-node workhorse.

Specifications

GLM-5.3-Flash at a glance.

SpecGLM-5.3-Flash
Architecture320B-parameter MoE · 18B active · new base (not GLM-5.2) · hybrid sparse + linear attention · Manifold-Constrained Hyper-Connections
ModalitiesNatively multimodal — text, image, video, visual documents (30T-token multimodal pre-training)
Context window1,000,000 tokens · 128K max output (evaluations published to 300K)
LicenseMIT — plain, unmodified, no regional restrictions
ReleasedAugust 25–26, 2026 · zai-org/GLM-5.3-Flash · GGUF, NVFP4 and Unsloth quants available
Benchmarks (model card)Terminal-Bench 2.1 84.3 · DeepSWE v1.1 63.4 (GLM-5.2: 46.2) · ExtractBench Short 96.3 · AutomationBench 48.8
Checkpoint sizesBF16 642 GB · Q8_0 341 GB · FP8 ≈ 306 GiB · UD-Q4_K_XL 200 GB · UD-Q2_K_XL 109 GB · UD-IQ1_S 93 GB (Unsloth GGUF)
ServingvLLM · SGLang · TokenSpeed · Transformers · KTransformers · llama.cpp / Ollama / LM Studio via GGUF

The hybrid attention is the engineering story: linear-attention layers interleaved with sparse layers cut the KV-cache cost of long contexts sharply, which is why a 1M-token window is practical on a single node here when it was not on GLM-5.2.

Hardware requirements

Sizing the deployment honestly.

Capabilities

What it’s best at.

GLM-5.3-Flash is the general-purpose model for regulated organizations in 2026: strong agentic coding and tool use (DeepSWE 63.4 against GLM-5.2’s 46.2), structured extraction (ExtractBench Short 96.3), and — for the first time in the GLM-5 series — native document, image and video understanding. That collapses the old two-model pattern: the workloads that needed GLM-5.2 plus a separate vision model now run on one checkpoint behind one gateway.

Z.ai prices the hosted version at roughly one-tenth of the full GLM-5.3; on owned hardware the gap shows up as memory — 200 GB at 4-bit against 245–400 GB — and as decode speed, since 18B active parameters run faster than ~40B. Where the last few points of coding autonomy matter, we deploy the full GLM-5.3; everywhere else, Flash.

Deployment patterns

How we deploy it.

The single-node sovereign workhorse. One hardened node on your premises or in in-country colocation, serving GLM-5.3-Flash with SSO, role-based access, and full audit logging. Documents, scans, and screenshots go straight to the model — no separate OCR stage. Four to eight weeks from assessment to production.

Fine-tuned house model. LoRA adapters trained on your precedents, templates, or archive, inside your environment. The MIT license keeps every derivative unambiguously yours — no revenue thresholds, no review clauses, nothing to renegotiate.

Agent runtime model. Pinned as the model behind self-hosted agent frameworks (OpenClaw 2.0, Hermes Agent) so the agents that read your inbox, files, and tickets never send a token off-site.

Questions we get

Frequently asked questions

What hardware does GLM-5.3-Flash need?

It depends on precision. The Unsloth 4-bit GGUF is about 200 GB, the 2-bit builds 93–109 GB, FP8 about 306 GiB, and BF16 642 GB, before KV cache. A two-GPU workstation with 96–141 GB accelerators or a four-GPU 80 GB node serves the 4-bit build to a workgroup; a single 8-GPU node serves FP8 firm-wide. Because only 18B of 320B parameters activate per token, throughput per GPU is high and the 1M-token context is genuinely usable.

Is GLM-5.3-Flash really MIT-licensed when GLM-5.3 is not?

Yes. GLM-5.3-Flash ships under the plain MIT license with no regional restrictions and no added clauses. The full GLM-5.3 uses a custom “GLM-5.3 License” that adds a security-review requirement for model-as-a-service operators above US$10B in revenue. For most enterprises that clause never applies, but Flash removes the question entirely.

Does GLM-5.3-Flash replace GLM-5.2?

For most deployments, yes. It beats GLM-5.2 across the published benchmarks, adds native vision, runs in less than half the memory, and keeps the MIT license GLM-5.2 had. The exception is a pure coding or agent-autonomy workload where the full GLM-5.3’s extra points justify the larger footprint. Existing GLM-5.2 racks re-provision to Flash with headroom to spare.

How does it compare with Qwen3.8-Flash-Next?

They are the two efficiency flagships of late August 2026. Qwen3.8-Flash-Next is smaller — 125B total with 6B active plus a 51B N-gram table that can live in system RAM — and runs in roughly 75–110 GB quantized, but it ships under the Qwen Community License rather than MIT and its native context is 262K (1M with YaRN). GLM-5.3-Flash is larger, MIT-licensed, 1M-context natively, and posts higher agentic-coding scores. We benchmark both on your documents before choosing.

Frontier-class, MIT-licensed, one node. Size yours.

The two-week sovereignty assessment sizes the hardware against your real workloads and hands you a written architecture with a cost model — before you buy a single GPU.

Book a sovereignty assessment