Model bench · DeepSeek

Deploy DeepSeek on-premise.

DeepSeek V4 is the price-performance leader of the open-weight field, and since August 2026 it is two MIT-licensed models: V4-Pro-0813, a 1.6–1.7-trillion-parameter mixture-of-experts with 49B active parameters and a 1M-token context that posts Terminal-Bench 2.1 87.9, and V4-Flash, a 284B/13B-active sibling whose Vision-Exp checkpoint (weights August 31) adds image understanding. On owned hardware they are the batch engine — the models that chew through document queues, analytics, and classification overnight at zero marginal cost.

Specifications

DeepSeek V4 at a glance.

SpecDeepSeek V4
V4-Pro-08131.6–1.7T MoE · 49B active · 1,048,576 context · 384K max output · ~893 GB checkpoint · deepseek-ai/DeepSeek-V4-Pro-0813
V4-Flash-Vision-Exp304.6B total (284B V4-Flash + vision tower) · 13B active · 1M context · ~168 GB (FP8/FP4 experts) · weights Aug 31, 2026
Benchmarks (model cards)Pro: Terminal-Bench 2.1 87.9 · DeepSWE 62.7 · HLE w/ tools 60.0. Flash-Vision: Terminal-Bench 2.1 83.9 · DeepSWE 59.3
ModalitiesText (Pro) · text + images at up to 384 tokens per image (Flash-Vision)
LicenseMIT — both models
ServingvLLM (dedicated V4 images) · SGLang · reasoning_effort low/high/max · DSpark speculative decoding on Pro
ProfileCost-efficient reasoning at scale; strongest price-performance for high-volume work
Serving classFlash: single 4-GPU node. Pro: 8-GPU node to small cluster, sized by throughput target

DeepSeek’s lineage defined the economics of open-weight reasoning: frontier-adjacent quality at a fraction of proprietary inference cost. On owned hardware, that efficiency compounds — throughput per dollar is the design goal.

Hardware requirements

Sizing the deployment honestly.

Capabilities

What it’s best at.

DeepSeek V4 is the model we reach for when the workload is measured in millions of tokens per day: batch document analysis, high-volume summarization and extraction, analytics over internal records, and reasoning-heavy classification. In the financial-services pattern it runs overnight batch analytics beside a GLM-5.2 member-service copilot; in legal and insurance work it powers first-pass review queues.

The structural advantage is economic. Metered APIs punish exactly this kind of adoption — every additional document costs more. Owned DeepSeek inverts the curve: per-token cost falls as volume grows, and the overnight queue runs free at the margin.

Deployment patterns

How we deploy it.

The batch engine. Deployed beside the interactive workhorse: GLM-5.2 answers people in real time, DeepSeek processes queues at night. One gateway routes each job to the cheapest model that clears your quality bar.

Pinned and documented. Like every model on our bench, DeepSeek deploys with pinned, checksummed weights, an evaluation harness, and version documentation — the model-risk file regulators now expect. Air-gapped operation is available.

Questions we get

Frequently asked questions

What hardware does DeepSeek V4 need on-premise?

A quantized deployment serving a workgroup runs on a multi-GPU inference node; full-precision batch clusters scale with your throughput target rather than your headcount. Because DeepSeek is efficiency-optimized MoE, tokens per dollar is the sizing variable — the sovereignty assessment converts your projected daily volume into a concrete bill of materials.

When does DeepSeek beat GLM-5.3 or Kimi K3 as the choice?

When volume dominates, or when the license must be MIT. For interactive copilots GLM-5.3-Flash is usually the better workhorse, GLM-5.3 leads on agentic coding, and K3 leads on deep autonomous work. DeepSeek V4-Pro-0813 is within a point of them on Terminal-Bench (87.9) under a plain MIT license, and V4-Flash-Vision brings MIT-licensed vision in about 168 GB. DeepSeek wins on high-volume batch reasoning — analysis, extraction, classification, summarization at millions of tokens per day — where its price-performance is the strongest in the open-weight field.

Is a Chinese-origin model a problem for our governance?

Self-hosted open weights are static files: audited, checksummed, incapable of transmitting anything, and runnable fully air-gapped. No data flows back to the lab that trained them. Where procurement policy nonetheless requires Western-origin models, we deploy Llama and Mistral instead — model origin is a governance choice, and sovereignty means you make it.

Turn the overnight queue into free compute.

The two-week sovereignty assessment sizes the hardware against your real workloads and hands you a written architecture with a cost model — before you buy a single GPU.

Book a sovereignty assessment