Qwen3.8-Flash-Next Hardware Requirements for Self-Hosting
Qwen3.8-Flash-Next self-hosting: 125B MoE, 6B active, FP8 at 173 GiB, 4-bit GGUF at 111 GB, N-gram table in system RAM, and the Qwen Community License.
Field notes
In-depth guides on on-premise LLM deployment: hardware requirements, cost economics, and the compliance landscape for regulated organizations.
Qwen3.8-Flash-Next self-hosting: 125B MoE, 6B active, FP8 at 173 GiB, 4-bit GGUF at 111 GB, N-gram table in system RAM, and the Qwen Community License.
OpenClaw 2.0 (v2026.8.1) vs Hermes Agent v0.21.0: local-model support, SQLite migration, Bot Mode, sandboxing, and how to run either fully on-prem.
GLM-5.3 on-prem guide: 753B MoE, ~40B active, 1M context, Terminal-Bench 2.1 88.2. Memory by precision, node layouts, and the new license clause explained.
GLM-5.3-Flash sizing guide: 320B MoE, 18B active, multimodal, 1M context, MIT license. Quant-by-quant memory table (93–642 GB), node layouts, what it replaces.
September 2026 comparison: GLM-5.3-Flash, GLM-5.3, Kimi K3, DeepSeek V4, Hy4, Qwen3.8, Granite 4.2, Muse Glimmer. Hardware, licensing, which to deploy.
Kimi K3 self-hosting guide: 2.8T MoE, 1M-token context, ~1.4TB native-4-bit checkpoint, deployment tiers, and Moonshot's 64+ accelerator guidance.
OSFI Guideline E-23 puts AI models under full model-risk management by May 1, 2027. What it requires — and why self-hosted models make the file defensible.
Deploy an LLM with zero network connectivity: CMMC and ITAR requirements, the architecture, and the telemetry and license-check traps that break isolation.
How Quebec Law 25 applies to AI tools: cross-border assessments, automated-decision transparency, and penalties up to C$25M or 4% of worldwide turnover.
GLM-5.2 deployment guide: 744B MoE with 40B active parameters, MIT license, 1M context — precision tiers, GPU counts, and single-rack serving that works.
Five architectures for using LLMs with PHI under HIPAA — and why a fully on-premise deployment needs no business associate agreement at all.
The US CLOUD Act (18 USC 2713) reaches data held by US providers even in Canadian regions. What that means for Canadian AI deployments — and the fix.
What on-premise LLM deployment actually costs in 2026: hardware tiers from a single node to a Kimi K3 cluster, and the ~2M tokens/day break-even vs APIs.
US v. Heppner (SDNY, Feb 2026) held consumer-AI chats are not privileged. What lawyers in Canada and the US can defensibly do with client files and AI.
Two weeks, fixed fee. We map your data obligations and workloads, size the hardware, and hand you a written architecture with a real cost model — whether or not you build with us.
Book a sovereignty assessment