Legal · Privilege
Frontier-class AI that never sees the other side of your firewall.
On-premise open-weight models give your lawyers discovery triage, precedent-aware drafting, and research over the firm’s entire work product — with privileged material never transiting a third party. Compliance with law society confidentiality duties and Québec Law 25the duty of confidentiality and attorney–client privilege is satisfied by architecture, not by vendor contract.
The problem
Heppner ended the ambiguity.
In US v. Heppner (SDNY, February 2026), Judge Rakoff held that conversations with consumer AI tools are not privileged. What a client — or a lawyer — types into a consumer chatbot is discoverable, like any other communication routed through a third party with its own logs and retention policies. The decision did not create the risk; it named a risk that was always there.
For Canadian firms, the exposure stacks. Law society rules of professional conduct impose a duty of confidentiality that covers every technology the firm adopts. Québec Law 25 requires an assessment before personal information crosses the border — and most frontier AI APIs are US-operated. The US CLOUD Act then lets American authorities compel those providers to disclose data they hold, even data sitting in Canadian cloud regions. A firm cannot contract its way around a foreign statute.
For US firms, ABA Model Rule 1.6 and its state analogues require reasonable efforts to prevent unauthorized disclosure of client information — and opinion after opinion now treats consumer AI tools as exactly the kind of third-party disclosure the rule contemplates. Every vendor in the chain adds logs, retention policies, and legal exposure that become the firm’s problem in discovery.
The result at most firms is the worst of both worlds: an official ban that associates quietly route around with personal accounts, because the tools are genuinely useful. Roughly 27% of employees admit to pasting confidential material into public AI tools. Bans don’t hold. A sanctioned tool that is actually better does.
The architecture
Privilege protected by physics, not paperwork.
We deploy flagship open-weight models — GLM-5.2 as the workhorse, Kimi K3 for cluster-scale work — on hardware the firm owns, inside the firm’s network. The weights are static files: downloaded, checksummed, and incapable of transmitting anything. There is no third-party API call in the chain, no vendor to subpoena, and no usage meter punishing adoption.
Fine-tuned on two decades of firm precedents, the model drafts in your house style, triages discovery across million-token document sets overnight at zero marginal cost, and answers research questions against the firm’s own work product. SSO, matter-scoped role-based access, and full audit logging carry your ethical walls into the AI layer.
Deployment blueprint
The 120-lawyer firm, on paper.
A TorontoChicago firm deploys GLM-5.2 LoRA-tuned on its precedent corpus: matter summarization, first-pass privileged discovery triage, and precedent-aware drafting — all inside the firm’s own network. Read the full reference architecture.
Read the law firm blueprintQuestions we get
Frequently asked questions
Can lawyers use ChatGPT or other cloud AI tools with client files?
Not defensibly. In US v. Heppner (SDNY, February 2026), Judge Rakoff held that consumer-AI conversations are not privileged — material a client or lawyer puts through a consumer chatbot is discoverable. Layer on law society confidentiality duties in Canada, Quebec Law 25 cross-border transfer assessments, the US CLOUD Act’s reach over American providers, and the duty of confidentiality under ABA Model Rule 1.6 in the US, and every privileged prompt to a third-party cloud carries stacked risk. An on-premise deployment removes the question: privileged material never leaves infrastructure the firm owns.
What can an on-premise model actually do for a law practice?
First-pass discovery triage across million-token document sets, matter summarization, precedent-aware drafting, deposition and transcript analysis, and internal research over the firm’s own work product. With a model fine-tuned on the firm’s precedents, output arrives in the firm’s formats and house style — and overnight batch review runs at zero marginal cost on hardware you already own.
Is an open-weight model good enough for legal work?
Yes. As of September 2026 the leading open-weight models — GLM-5.3, GLM-5.3-Flash and Kimi K3 — perform at the level of last-generation proprietary flagships on reasoning, drafting, and long-document benchmarks, and handle 1M-token contexts, enough for entire document productions in a single pass. Kimi K3 is also the best open model at admitting uncertainty rather than hallucinating, which matters in work where a fabricated citation is a professional-conduct problem.
How is privilege protected in the deployment itself?
The model runs inside the firm’s network with SSO, role-based access scoped by matter, and full audit logging. No third-party API calls, no vendor logs, no vendor retention policy — the firm remains the only custodian of privileged material, and ethical walls carry through to the AI layer.
What does deployment cost for a mid-sized firm?
A single-rack deployment serving GLM-5.2-class models (40B active parameters) to a firm-wide workgroup starts in the low-to-mid six figures of hardware, with break-even against metered APIs around two million tokens per day of usage. The sovereignty assessment produces a real cost model against your projected usage before any commitment.
Go deeper
Can lawyers use ChatGPT with client files?
What US v. Heppner changed, and what a defensible posture looks like.
The CLOUD Act and Canadian data
Why “hosted in Canada” is not enough when the provider is American.
Québec Law 25 and AI
The “equivalent protection” bar and penalties to $25M / 4% of turnover.
Deploy GLM-5.2 on-premise
The single-rack workhorse: 744B MoE, 40B active, MIT-licensed.
Deploy Kimi K3 on-premise
The largest open model released — 2.8T parameters, 1M-token context.
Fine-tuning on firm precedents
LoRA adapters trained inside your environment. Your corpus never leaves.
Put privileged material back behind your own walls.
The two-week sovereignty assessment maps your confidentiality obligations, matters, and workloads — and hands you a written architecture and cost model, whether or not you build with us.
Book a sovereignty assessment