Model bench · Moonshot AI

Deploy Kimi K3 on-premise.

Kimi K3 is the largest open-weight model ever released: a 2.8-trillion-parameter mixture-of-experts (104B active) with a 1M-token context window and native vision, which beat Claude Opus 4.8 and GPT-5.5 on coding and agentic benchmarks at launch. Its open weights have been on Hugging Face since July 27, 2026 — a 1.56 TB checkpoint you can download, checksum, and serve entirely inside your perimeter. The question is no longer access; it is whether your workloads justify the hardware.

Specifications

Kimi K3 at a glance.

SpecKimi K3
Architecture2.8T-parameter mixture-of-experts (MoE), 104B active — 16 of 896 experts + 2 shared per token; Kimi Delta Attention + Gated MLA; MoonViT-V2 vision encoder
Context window1,000,000 tokens
ModalitiesText + native vision (screenshots, documents, UI, images)
Published weights≈1.56 TB checkpoint · moonshotai/Kimi-K3 · MXFP4 weights supported
Weights availabilityOpen weights since July 27, 2026 · Kimi K3 License (custom — read before commercial redistribution)
Serving class (full precision)Supernode-class: Moonshot recommends 64+ accelerators
Benchmark positionGPQA Diamond 93.5 · Terminal-Bench 2.1 88.3 · DeepSWE 67.5 (model card). Beat Claude Opus 4.8 and GPT-5.5 on coding & agentic benchmarks at launch
Distinguishing traitBest open model at admitting uncertainty rather than hallucinating

Independent reviewers verified multi-hour autonomous runs — 3+ hours and 122 tasks completed from a single paragraph-long prompt — plus sub-agent orchestration and vision-in-the-loop front-end work in the first 24 hours after launch.

Hardware requirements

Sizing the deployment honestly.

Capabilities

What it’s best at.

K3’s profile is agentic depth. It sustains multi-hour autonomous sessions — codebase migrations, overhauls, and 100+ task runs from a single prompt — and orchestrates swarms of parallel sub-agents that divide work, verify each other, and synthesize results. Native vision closes the loop: it screenshots the UI, dashboard, or document it just produced, critiques it, and iterates.

On desk work it leads spreadsheet and knowledge-work benchmarks: research reports, financial analysis, and document review across million-token contexts. And for regulated work, its most important trait is negative: it is the best open model at saying “I don’t know” instead of fabricating — which matters when a hallucinated citation is a professional-conduct problem.

Deployment patterns

How we deploy it.

Pinned, checksummed, served. The checkpoint is downloaded once, verified against the published hashes, and served inside your perimeter behind an evaluation harness — with no interim-API exposure of sensitive data at any point.

Paired-bench deployment. The common pattern pairs K3 (cluster tier, deep reasoning and agentic work) with GLM-5.3-Flash (single-node, high-volume daily inference) behind one gateway — each request routed to the cheapest model that clears your quality bar. Air-gapped operation is available at every tier.

Questions we get

Frequently asked questions

What hardware does Kimi K3 require?

Full-precision serving is supernode-class: Moonshot recommends 64 or more accelerators, which puts it in large enterprise cluster or managed sovereign facility territory. Expert-offload and MXFP4 configurations bring evaluation and batch workloads within reach of large single nodes. The published checkpoint is roughly 1.56 TB, and sizing depends on offload strategy, context length in active use, and concurrency — which is exactly what the sovereignty assessment quantifies against your workloads.

Are the Kimi K3 weights actually available?

Yes. Moonshot published the checkpoint on Hugging Face (moonshotai/Kimi-K3) on July 27, 2026 under the Kimi K3 License. Before that date the only access was Moonshot’s own API, routed through infrastructure in China — acceptable for evaluation with non-sensitive material and nothing more. Today the defensible path for client files, patient records, or proprietary code is the downloaded, checksummed checkpoint served on hardware you own.

Is Kimi K3 actually competitive with GPT and Claude?

At launch, K3 beat Claude Opus 4.8 and GPT-5.5 on coding and agentic benchmarks and ranked third on independent intelligence indexes, behind only the two absolute proprietary flagships. The newest proprietary models still lead some evaluations — but they are also the models whose foreign access was cut off by export order in June 2026. Owned flagship-tier intelligence, available 24/7 with no usage caps, outperforms rented access you can lose overnight.

Do most organizations actually need K3, or is GLM-5.3-Flash enough?

Most workloads run superbly on GLM-5.3-Flash, whose 18B-active design puts a frontier-class, multimodal, MIT-licensed model on a single node (about 200 GB at 4-bit). K3 earns its cluster when the workload is deep agentic autonomy, vision-in-the-loop work, or maximum reasoning quality on high-stakes documents. We benchmark both against your real tasks and tell you honestly which tier your workloads justify.

The weights are public. Size the cluster.

The two-week sovereignty assessment sizes the hardware against your real workloads and hands you a written architecture with a cost model — before you buy a single GPU.

Book a sovereignty assessment