The model bench

Flagship open-weight models, deployed as of September 2026.

Five frontier-class open-weight releases landed in nine days at the end of August 2026 — GLM-5.3-Flash, Qwen3.8-Flash-Next, GLM-5.3, Tencent Hy4 and DeepSeek V4-Flash-Vision. These are the models we deploy today — selected, quantized, and served to match your workload and governance requirements, on infrastructure you own.

ModelScaleContextStrengthsLicense / weights
GLM-5.3-Flash
Zhipu AI / Z.ai
320B MoE · 18B active — 4-bit fits in ~200 GB; the new single-node workhorse1M tokens · native visionGeneral assistant, agentic coding and tool use, document/image/video understanding — one checkpoint replaces the old GLM-5.2 + vision-model pair. Deployment guide →MIT · Aug 25–26, 2026
GLM-5.3
Zhipu AI / Z.ai
753B MoE · ~40B active — in-place upgrade of GLM-5.21M tokens · text-onlyThe strongest open agentic-coding model: Terminal-Bench 2.1 88.2, DeepSWE 66.9, security research. For engineering organizations. Deployment guide →GLM-5.3 License · Aug 28, 2026
Kimi K3
Moonshot AI
2.8T MoE · 104B active — the largest open model released1M tokens · native visionDeep multi-hour autonomous runs, vision-in-the-loop work, document & spreadsheet analysis. Supernode-class hardware. Deployment guide →Kimi K3 License · weights Jul 27, 2026
DeepSeek V4
DeepSeek
V4-Pro-0813: 1.6–1.7T MoE · 49B active · V4-Flash: 284B · 13B active1M tokens · Flash-Vision variant adds imagesCost-efficient frontier reasoning at scale; the batch engine for high-volume analysis, extraction and classification. Deployment guide →MIT
Qwen3.8 family
Alibaba
Flash-Next: 125B · 6B active (+51B N-gram table in RAM) · 27B dense · 2.4T Max262K native → 1M · multimodalEfficiency and document/vision work: Flash-Next runs in ~75–110 GB quantized; Qwen3.8-27B is the Apache-2.0 workstation model. Deployment guide →Qwen Community 1.0 / Apache 2.0 (27B)
Llama · Mistral
Meta / Mistral AI
Multiple scalesLong contextWestern-origin open weights for organizations whose procurement or security policy requires them. Deployment guide →Open weights

On model origin: self-hosted open weights are static files — audited, checksummed, and incapable of transmitting anything. No data flows back to the lab that trained them, and fully air-gapped operation is available. Where policy requires Western-origin models regardless, we deploy them. Sovereignty means the choice is yours.

On honest sizing: full-precision Kimi K3 is supercomputer-class — Moonshot itself recommends 64+ accelerators. Most organizations don’t need that. GLM-5.3-Flash’s 18B-active design puts a frontier-class, multimodal, MIT-licensed model on a single node (about 200 GB at 4-bit); the full GLM-5.3 fills an 8-GPU node; K3 is available through large-cluster builds or our managed sovereign facility tier. The assessment tells you which tier your workloads actually justify.

On licenses: licensing is now the differentiator. GLM-5.3-Flash, DeepSeek V4 and Mistral Large 3 are MIT or Apache 2.0; Tencent Hy4, IBM Granite 4.2, Meta Muse Glimmer and Qwen3.8-27B are Apache 2.0; GLM-5.3, Kimi K3, Qwen3.8-Max and Qwen3.8-Flash-Next ship under custom licenses that are fine for internal enterprise use but need a procurement read. Every deployment file we hand over includes the license analysis.

Start with the sovereignty assessment.

Two weeks, fixed fee. We map your data obligations and workloads, size the hardware, and hand you a written architecture with a real cost model — whether or not you build with us.

Book a sovereignty assessment