Solution · Route

Frontier where it earns its price. Open weights where the data lives. One gateway in front of both.

We deploy Claude, GPT-5.6, Gemini and Grok beside open and customized weights on your own hardware, through the agent harnesses your teams actually use — Hermes Agent, OpenClaw, Claude Code, Grok Bot — behind a routing layer you control. Every call is classified, budgeted, redacted where needed, logged, and has a fallback.

The pattern

Three tiers, routed per task.

TierWhat runs thereWhat it seesExamples
FrontierThe hardest reasoning, novel synthesis, long-horizon planning where the best model measurably winsDe-identified or public inputs only, through region-pinned endpoints (Bedrock, Vertex, Azure Data Zones) with no-training termsClaude Fable 5.1 / Opus 5 / Sonnet 5 · GPT-5.6 Sol / Terra / Luna · Gemini 3 Pro · Grok 4.6
Open weightsVolume, latency-sensitive and everything that touches client, patient, deal or project dataEverything — it never leaves your hardwareGLM-5.3-Flash · GLM-5.3 · DeepSeek V4 · Qwen3.8 · Kimi K3
Customized weightsHouse style, formats, cost codes, terminology — the repetitive institutional work where consistency is the productYour corpus, trained inside your environmentLoRA adapters on the open-weight tier, evaluated on your harness — see fine-tuning

Why now: the curves crossed in August 2026. GLM-5.3 posts Terminal-Bench 2.1 88.2 against Kimi K3’s 88.3 from downloadable weights, and GLM-5.3-Flash fits a single node under MIT. The frontier tier still leads on the hardest problems and is priced for it — Fable 5.1 at $10 in / $50 out per million tokens, Sonnet 5 at $2 / $10, GPT-5.6 from $0.20 (Luna) to $10 (Sol, long context) input. Routing per task is what turns that gap into savings rather than a standardization debate.

The gateway

One door for every model, owned by you.

Self-hosted. We deploy LiteLLM, Bifrost or Kong AI Gateway inside your perimeter — never a hosted router that sees your prompts. Per-team and per-project hard budgets, an audit record for every request, retries and fallbacks, semantic caching, and PII redaction before anything reaches a frontier endpoint. Bifrost is also an MCP gateway, so tool permissions are policed in the same place.

Per-step routing. The next step past per-request routing is routing each step of an agent workflow to the model that fits it — NVIDIA’s open NeMo Switchyard library is the reference, reporting task cost cut to a third in its own tests. We design agent workflows so the planning step, the drafting step and the verification step can each land on a different tier.

Fallbacks that were tested. In June 2026 Anthropic suspended Fable 5 and Mythos 5 for every customer for almost three weeks under a US export order. Every architecture we ship has an open-weight fallback for each frontier route, exercised in a drill before go-live, so a vendor event degrades capability rather than stopping work.

Serving under the gateway. On the open-weight side we run vLLM or SGLang, and NVIDIA Dynamo (1.0, March 2026) where disaggregated prefill and decode across a rack pays for itself.

Harnesses

The runtimes your agents live in.

A harness is where agents get their permissions, memory, schedule and channels. We deploy the ones with real permission, audit and MCP stories, and integrate the ones your teams already pay for.

HarnessWhat it isWhere we use itBlueprint
Hermes Agent (Nous Research, MIT)v0.21 “Pantheon”: bots as profiles with a pinned model each, agent inbox, bot-to-bot messaging, scheduled runs with memory, MCP command centre, approval before editing skills or memory; gateways to Telegram, Slack, WhatsApp, SignalFleets of role-based bots for research, drafting, inbox, scheduling and compliance checks over a shared source of truthHermes fleet with hybrid routing
OpenClaw 2.0 (open source, nonprofit foundation)v2026.8.1: guided install that detects local models, in-process GGUF and managed llama-server, per-operation permissions, sandboxing, masked credentials, 1Password broker, SQLite sessionsAlways-on personal agents on workstations and a shared server, with local models by defaultOpenClaw sovereign workstations
xAI Grok BotAlways-on agents with their own cloud computer, signed into a customer’s tools; bundled with Cursor Ultra and SuperGrok Heavy; enterprise waitlist; Grok is also a provider inside Hermes AgentLow-sensitivity work for teams that already have it, bridged to a self-hosted fleet for anything with client dataGrok Bot beside a Hermes fleet
Claude Code · Managed Agents (Anthropic)Terminal and IDE coding agent; Managed Agents GA April 2026 for server-hosted agents with a managed sandbox; MCP nativeEngineering organizations on the frontier tier, with repository access scoped by the gateway
OpenAI Agents SDK · Google ADK 2.0 · Microsoft Agent Framework 1.0Vendor SDKs, all GA in 2026, all with MCP supportWhere a team’s stack is already committed to one cloud; routed through the same gateway

Engagement

From model sprawl to a governed stack.

Questions we get

Frequently asked questions

Can we use frontier models like Claude or GPT-5.6 and still be sovereign?

For the right workloads, yes — with three controls. Region-pinned processing: Claude runs on AWS Bedrock and Google Vertex infrastructure in-region, and OpenAI offers Azure Data Zones; we route only to endpoints that process in your jurisdiction. Data classification: anything that identifies a client, patient or deal goes to open weights on your own hardware, and the frontier tier sees de-identified or public inputs only. And a gateway you own in front of every call, so budgets, redaction, fallbacks and the audit log are uniform whichever model answers.

Why not pick one model and standardize?

Because the price and capability curves crossed. Open weights like GLM-5.3 and Kimi K3 now post frontier-class agentic-coding scores, and GLM-5.3-Flash fits one node under an MIT license. Frontier models still lead on the hardest reasoning and are priced accordingly — Claude Fable 5.1 at $10 per million input tokens and $50 output, Sonnet 5 at $2 and $10, GPT-5.6 in three tiers from $0.20 to $10 input. Routing per task keeps the frontier bill for the work that needs it and runs the rest at owned-hardware cost. NVIDIA reports that per-step routing cut task cost to a third in its own tests.

What happens when a frontier vendor cuts access or changes terms?

June 2026 is the reference case: a US export order forced Anthropic to suspend Fable 5 and Mythos 5 for every customer worldwide for almost three weeks, because there was no way to gate by nationality in real time. Workflows with a single frontier dependency stopped. Workflows behind a gateway with an open-weight fallback kept running at reduced capability. That is the design goal of every hybrid architecture we build.

Which agent harnesses do you deploy?

The ones with real permission, audit and MCP stories: Hermes Agent (profiles with pinned models, an agent inbox, scheduled runs with memory, approval before editing skills or memory), OpenClaw 2.0 (per-operation permissions, sandboxing, local GGUF inference), Claude Code and Managed Agents, the OpenAI Agents SDK, Google ADK and the Microsoft Agent Framework. We also integrate xAI Grok Bot where a team already has it, noting that it currently ships bundled with Cursor Ultra or SuperGrok Heavy with no enterprise plan. Every harness sits behind the same gateway and the same source of truth.

Can a frontier model run on our premises?

Google is the only frontier vendor with a shipping on-premise appliance — Gemini on Google Distributed Cloud, Dell-built on NVIDIA Blackwell, air-gappable. OpenAI announced a Dell partnership for on-premise Codex in May 2026 with pricing undisclosed. Anthropic and xAI have no on-premise offering. Where the workload demands both frontier capability and no egress, the answer today is open weights on your hardware, which is why the bench matters.

Route the work, not the standardization debate.

The two-week sovereignty assessment maps where every prompt goes today, and hands you a routing design with a gateway, a fallback plan and a cost model — before you sign another model contract.

Book a sovereignty assessment