Model bench · Zhipu AI / Z.ai
Deploy GLM-5.2 on-premise.
GLM-5.2 was the practical on-premise workhorse of mid-2026: a 744-billion-parameter mixture-of-experts with only 40B active parameters per token, a 1M-token context window, and a fully permissive MIT license. Released June 13, 2026 with no regional restrictions, it landed within a few points of Claude Opus 4.8 on key agentic benchmarks. As of September 2026 it is superseded: the same base model ships as GLM-5.3 with far stronger coding results, and the new GLM-5.3-Flash beats it across the published benchmarks in under half the memory while keeping the MIT license. Existing GLM-5.2 racks upgrade to GLM-5.3 by swapping checkpoints; most re-provision to Flash. This page remains as the sizing reference for deployments already running it.
Specifications
GLM-5.2 at a glance.
| Spec | GLM-5.2 |
|---|---|
| Architecture | 744B-parameter mixture-of-experts (MoE), 40B active parameters per token |
| Context window | 1,000,000 tokens |
| Modalities | Text-only — pair with Qwen3-VL for vision workloads |
| License | MIT — fully permissive, no regional restrictions |
| Released | June 13, 2026 · open weights immediately |
| Benchmark position | Within a few points of Claude Opus 4.8 on key agentic benchmarks, at roughly one-fifth the cost |
| Serving class | Single-rack on-premise deployment — the 40B-active design is the enabler |
The MIT license matters as much as the benchmarks: no usage restrictions, no acceptable-use gatekeeping, no license terms to renegotiate. The model is yours to run, fine-tune, and keep.
Hardware requirements
Sizing the deployment honestly.
Quantized workgroup tier
Quantized GLM-5.2 serves a professional workgroup — a practice group, a clinic network, a department — from a single hardened inference node. This is the entry point for most organizations: low-to-mid six figures of hardware, full sovereignty properties, room to grow.
Full-precision rack tier
Full-precision serving with headroom for 1M-token contexts and firm-wide concurrency runs on a single GPU rack. Per-token cost falls as usage grows, and overnight batch queues — document review, analysis, drafting — run at zero marginal cost. Break-even against metered APIs arrives around 2M tokens/day.
Capabilities
What it’s best at.
GLM-5.2’s core strengths are software engineering and tool-driven agents: repository-scale coding, structured tool use, and reliable multi-step task execution. It is the model we deploy most — the copilot behind law-firm drafting, credit-union member service, and manufacturing knowledge systems — because it delivers near-flagship agentic quality at a hardware footprint a mid-sized organization can actually own.
The 1M-token context carries entire document productions, codebases, or policy manuals in a single pass. For vision workloads — scanned records, drawings, forms — we pair it with Qwen3-VL behind the same gateway.
Deployment patterns
How we deploy it.
The single-rack sovereign node. The signature deployment: one hardened rack on your premises or in in-country colocation, serving fine-tuned GLM-5.2 with SSO, role-based access, and full audit logging. Six to twelve weeks from assessment to production.
Fine-tuned house model. LoRA adapters trained on your precedents, templates, or archive — in your environment — make GLM-5.2 the model that speaks your firm’s language. MIT licensing keeps every derivative unambiguously yours. Air-gapped operation is fully supported.
Questions we get
Frequently asked questions
What GPU hardware does GLM-5.2 require?
The 40B-active MoE design is the key: although the full parameter count is 744B, per-token compute is that of a 40B model, which puts serving within reach of a single GPU rack. Quantized workgroup deployments start in the low-to-mid six figures of hardware; full-precision, firm-wide serving with long-context headroom fills a rack. Exact GPU counts depend on precision, context usage, and concurrency — the sovereignty assessment produces the sized bill of materials.
Is the MIT license really unrestricted for commercial use?
Yes. MIT is the most permissive mainstream license in software: commercial use, modification, fine-tuning, and internal deployment are all unambiguously permitted, with no regional restrictions and no revocation mechanism. Your fine-tuned adapters and every derivative remain your property.
How does GLM-5.2 compare to Kimi K3?
K3 is deeper — stronger on long-horizon agentic autonomy and equipped with native vision — but full-precision K3 is supercomputer-class, needing 64+ accelerators. GLM-5.2 lands within a few points of last-generation proprietary flagships on agentic benchmarks while fitting on one rack. Most organizations run GLM-5.2 as the daily workhorse and add K3 at the cluster or managed-facility tier only where workloads demand it.
Go deeper
Deploy GLM-5.3 on-premise
The in-place upgrade: same base, Terminal-Bench 3.0 from 4.6 to 28.3.
Deploy GLM-5.3-Flash on-premise
MIT, multimodal, and under half the memory — where most GLM-5.2 racks are going.
GLM-5.2 GPU requirements
The full deployment guide: memory math, quantization, and rack design.
On-premise LLM deployment
The end-to-end engagement that gets GLM-5.2 racked and serving.
Deploy Qwen3-VL on-premise
The vision companion for scanned records, drawings, and forms.
The workhorse fits in one rack. Size yours.
The two-week sovereignty assessment sizes the hardware against your real workloads and hands you a written architecture with a cost model — before you buy a single GPU.
Book a sovereignty assessment