Solution · Build
Flagship open-weight models, running on servers you own.
On-premise LLM deployment is the complete build-out of private AI infrastructure: GPU cluster design and procurement, serving stack, security hardening, and production rollout of models like Kimi K3, GLM-5.2, and DeepSeek — with every token generated on hardware in Canada, under Canadian jurisdictionin the United States, under your control, and zero third-party API calls.
Who it’s for
Organizations whose best data can’t leave.
This engagement is built for regulated and IP-intensive organizations — law firms, healthcare networks, financial institutions, engineering and defense firms — where the data that would make AI most valuable is the data that cannot transit a third-party cloud. It also fits any organization whose usage has crossed the economic threshold: past roughly two million tokens per day, owned infrastructure beats metered APIs on cost alone, before a single compliance consideration.
The output is not a proof of concept. It is production infrastructure: SSO, role-based access, audit logging, observability, and a security review completed before any real data touches the system.
The engagement
From assessment to production in six to twelve weeks.
- Sovereignty assessment (weeks 1–2). We map your data classes, obligations, workloads, and real concurrency. Output: a written architecture and cost model against your projected usage.
- Architecture & hardware (weeks 2–6, parallel). GPU cluster design, procurement, and site preparation — premises, Canadian colocationdomestic colocation, or air-gapped.
- Deploy & harden (weeks 5–8). Serving stack, SSO, role-based access, audit logging, monitoring. Security review before any real data enters.
- Tune & evaluate (weeks 7–10). Fine-tuning on your corpus, in-environment, with an evaluation harness built from your actual work.
- Operate & upgrade (ongoing). Staged rollout, training, and the managed retainer: patches, model evaluations, and migrations when new releases win.
Deliverables
What you own at the end.
- A written sovereignty architecture and cost model, benchmarked against cloud spend.
- Production GPU infrastructure — sized, procured, racked, and hardened.
- A serving stack with SSO, role-based access, audit logging, and observability.
- Pinned, checksummed model weights with documented evaluation results.
- Fine-tuned adapters trained on your corpus — your property, portable across upgrades.
- Runbooks, admin training, and a staged rollout plan for your teams.
Questions we get
Frequently asked questions
What does on-premise LLM deployment cost?
It scales with model size, precision, and concurrency. A single-rack deployment serving GLM-5.2-class models (40B active parameters) to a workgroup starts in the low-to-mid six figures of hardware. Full-precision Kimi K3 is supercomputer-class — Moonshot recommends 64 or more accelerators — and is offered through large enterprise cluster builds or a managed sovereign facility tier. At every tier the cost curve differs structurally from cloud APIs: per-token cost falls as usage grows, break-even against metered APIs arrives around two million tokens per day, and overnight batch work runs at zero marginal cost.
Do we need our own data centre to run sovereign AI?
No. Three deployment patterns work: a hardened server room on your premises, your own rack in an in-country colocation facility, or managed hosting on hardware dedicated to you in a domestic facility. Data residency and control remain yours in every pattern — the difference is only who handles physical operations.
How long does deployment take?
Six to twelve weeks from kickoff to production in most engagements: sovereignty assessment and architecture in the first two weeks, hardware procurement in parallel, then deployment, hardening, fine-tuning, and staged rollout. Managed operations continue from there under the maintenance retainer.
Start with the sovereignty assessment.
Two weeks, fixed fee. We map your data obligations and workloads, size the hardware, and hand you a written architecture with a real cost model — whether or not you build with us.
Book a sovereignty assessment