Compute & AI Infrastructure
Kimi K3 vs DeepSeek V4 Flash: The Open-Weight Frontier
Open weights at frontier scale do not push capability to the edge, they move it to whoever can keep 1.4 terabytes loaded, and what that means for founders choosing what to self-host
Atomic answer
The open-weight frontier is a two-tier story. DeepSeek V4 Flash is an efficiency-first model that teams can deploy against, while Kimi K3 is frontier-grade and open but requires roughly 1.4 terabytes in four-bit format and on the order of eighteen 80 gigabyte accelerators to load. At that scale, open weights shift capability toward whoever can afford to keep them resident and busy.
Who is this for?
This article is for the technical founder or platform operator who is deciding, this quarter, which workloads to run on open weights and which to keep on a hosted API.
You are running an agent, coding assistant, document-processing, or retrieval workload. You have been told that open weights mean sovereignty, price control, and data residency. The decision in front of you is concrete: do you commit engineering time and capital expenditure to self-hosting a frontier open-weight model, or do you serve a smaller open-weight model at commodity prices and buy frontier capability from an API when a workload actually needs it?
It is also for the allocator underwriting AI infrastructure companies, and for the policy or export-control analyst tracking what open model weights actually enable at the margin. For all three readers the same question decides the answer: what does it cost to hold the weights, and what does the model actually do once you have them?
Where does the open-weight frontier actually bottleneck?
Not at the download. At the resident memory footprint.
Kimi K3's open weights shipped in 96 shards on Hugging Face on 2026-07-26 to 27. The full model requires about 1.4 terabytes of storage in MXFP4 four-bit format, and roughly 5.6 terabytes at 16-bit precision. Holding 1.4 terabytes of weights implies something on the order of eighteen 80 gigabyte accelerators just to load the model, which in practice means Nvidia Blackwell or AMD MI400 class hardware, and a single high-end server node can barely fit the weights with almost nothing to spare. (techi.com, 2026-07; qz.com on the weight drop.) Source confidence: Mixed (named public reporting on the release, combined with editorial interpretation of the deployment implication). Signal strength: High.
That is the whole constraint. A model you can legally download but cannot economically keep resident is not portable in any sense a builder can act on. The correct statement is close to the inverse of the usual open-weight narrative: at this scale, open weights move capability from model labs to infrastructure operators, and concentrate rather than diffuse deployment.
The architecture explains the footprint. Kimi K3 is reported at roughly 2.8 trillion total parameters with about 104 billion active per token (16 of 896 experts), a 1 million token context, and native vision. Those parameter and context claims originate with Moonshot and are now independently corroborated by multiple third parties. (Moonshot model card; corroborated in the joint institute assessment and in independent release coverage.) Source confidence: Mixed (originating source not opened this run; independently corroborated). Sparse activation cuts compute per token. It does not cut the memory you must hold.
DeepSeek V4 Flash is the other tier and it inverts every one of those numbers. It is an efficiency-optimized mixture-of-experts model with 284 billion total and 13 billion activated parameters, a 1 million token context, hybrid attention, and high and maximum reasoning effort levels, released 2026-04-24. It is listed at $0.0868 per million input tokens and $0.1736 per million output tokens, a 38 percent promotional discount, with prompt caching reducing the effective price further on repeated context. (OpenRouter model page, accessed 2026-08-02.) Source confidence: Primary (openrouter.ai, accessed 2026-08-02). Signal strength: High.
One caution on comparing the two: the price above is for a 13 billion active parameter model. Kimi K3 activates roughly 104 billion parameters per token and has no published comparable price. Those are different cost structures, and a per-token figure from one does not characterize the economics of the other. Cost, speed, and latency for Kimi K3 remain unverified against an independent profile as of 2026-08-02 and are not stated here.
Who controls the frontier once the weights are public?
Three actors, and none of them is the model lab alone.
The licensor. Kimi K3's weights shipped under a custom Kimi K3 License rather than a standard permissive license such as MIT or Apache 2.0. Commercial terms should be read in full before any hosting commitment; a non-standard license is a procurement question, not a footnote. Source confidence: Mixed (reported in independent release coverage; the license text itself should be read directly by counsel before commitment). Signal strength: High for the procurement implication.
The infrastructure operator. Whoever can keep 1.4 terabytes loaded and busy sets the terms on which everyone else gets frontier open-weight capability. That is a small population: hyperscalers, well-funded inference providers, and a handful of national or corporate compute programs. The practical effect of the release is to create a new hosting market, not a new class of self-hosters. Source confidence: Analytical (Stack & State ecosystem observation and pattern recognition). This is a hypothesis for navigation, not a verified finding.
The evaluator. On 2026-07-23 the UK AI Security Institute and the United States Center for AI Standards and Innovation published a joint preliminary assessment of Kimi K3's cyber capabilities. On "The Last Ones" attack path, Kimi K3 reached step 17 of 32 on average, against 28.5 for the most cyber-capable United States closed models. On ExploitBench it scored 32 percent against 24 percent for GLM-5.2, the open-weight model the institutes use as a comparison baseline. It achieved arbitrary code execution on 0 of 41 samples, against 20 of 41 for the most capable models on average. It completed the cyber range in 1 of 10 attempts within the token limit, and its safeguards did not prevent it from attempting cyber exploit development. (UK AI Security Institute, 2026-07-23.) Source confidence: Primary (aisi.gov.uk, accessed 2026-08-02). Signal strength: High.
Every one of those scores is specific to the institutes' own evaluation design and should not be generalized beyond it. Two further limits belong on the record. The assessment is explicitly preliminary, which means the findings are provisional. And an evaluation of a foreign laboratory's model conducted by two government institutes carries an institutional perspective that a reader should weigh alongside the numbers, not instead of them. Source confidence: Analytical.
Why should builders care about the intelligence-versus-cyber gap?
Because it is the only part of this story that tells you how to tier your workloads.
Kimi K3 outperforms the open-weight comparison baseline on ExploitBench while achieving arbitrary code execution on none of 41 samples. That pattern, better than other open weights and materially behind closed frontier models on the specific capability being measured, is what capability diffusion actually looks like in the middle of the curve. Benchmark convergence on general intelligence does not carry across to cyber capability, and it does not carry across to speed, cost, latency, or operational maturity at all.
Three practical consequences.
First, tier by consequence, not by benchmark. A workload where a wrong answer is cheap and reviewable is a good candidate for a self-hosted or commodity open-weight model. A workload with privileged system access or irreversible actions is not, regardless of how a model scores on a general benchmark.
Second, audit model card claims against a named assessment before making a hosting decision. The joint assessment gives deployers a citable, scope-limited reference for what an open-weight model does under adversarial evaluation. That is a better diligence artifact than a vendor benchmark table, provided its preliminary status is carried along with it.
Third, price the accelerators, not the tokens. For frontier open weights, the deployment decision is a capital expenditure and utilization question. Eighteen accelerators held at low utilization is worse economics than an API call, and the break-even is set by sustained throughput, not by the headline per-token price of a different and much smaller model.
What would change this assessment
- Portability. A credible quantization, offload, or serving technique that runs Kimi K3 class weights on a materially smaller accelerator footprint, with published throughput and quality numbers, would move the portability constraint from High to Medium. So would a published price for hosted Kimi K3 inference that lands near commodity levels.
- Deployment readiness. A full, non-preliminary assessment from the institutes, or an independent third-party evaluation that reproduces or contradicts the reported scores, would change the confidence attached to the cyber findings in either direction.
- License. Publication of clear, permissive commercial terms under the Kimi K3 License would remove a procurement constraint that currently sits between the weights and a hosting commitment.
- Cost basis. An independent cost, speed, and latency profile for Kimi K3, which was not verified this run, would allow the two-tier comparison to be made on like-for-like terms rather than structurally.
Last editorially reviewed: 2026-08-02.
FAQ
Q: Are Kimi K3's open weights actually usable by a startup? A: Not self-hosted, for almost all startups. About 1.4 terabytes of weights in four-bit format implies on the order of eighteen 80 gigabyte accelerators to load the model, and a single high-end server node can barely fit them (techi.com, 2026-07). The realistic path is renting the model from an operator who has already made that capital commitment. Source confidence: Mixed.
Q: Does Kimi K3 mean open weights have caught up with closed frontier models? A: Not on the capability that was actually measured. The joint UK AI Security Institute and United States Center for AI Standards and Innovation assessment reports Kimi K3 at step 17 of 32 on "The Last Ones" against 28.5 for the most cyber-capable United States closed models, and arbitrary code execution on 0 of 41 samples against 20 of 41 (UK AI Security Institute, 2026-07-23). It does outperform GLM-5.2, the institutes' open-weight baseline, on ExploitBench at 32 percent against 24 percent. All of these are specific to that evaluation design and the assessment is explicitly preliminary. Source confidence: Primary for the figures; Analytical for the read.
Q: What is the actual cost difference between the two models? A: It cannot be stated on a like-for-like basis today. DeepSeek V4 Flash is listed at $0.0868 per million input tokens and $0.1736 per million output tokens at a 38 percent promotional discount (OpenRouter, accessed 2026-08-02), but that is a 13 billion active parameter model. Kimi K3 activates roughly 104 billion parameters per token and has no published comparable price; its cost, speed, and latency profile remains unverified as of 2026-08-02. Treat the DeepSeek figure as the floor of the market, not as the price of frontier open weights. Source confidence: Primary for the DeepSeek pricing.
Q: Does the custom license matter if we are only running inference internally? A: It matters enough to read before you commit engineering time. The weights shipped under a custom Kimi K3 License rather than a standard permissive license, and non-standard terms can restrict commercial use, redistribution, or downstream service provision in ways that only surface at contract review. This is a procurement question for counsel, not a technical one. Nothing here is legal advice. Source confidence: Mixed.
Q: What should we self-host, then? A: Workloads where a smaller open-weight model is sufficient and where residency, price stability, or latency control justify the operational burden. DeepSeek V4 Flash's profile, 284 billion total and 13 billion active parameters with a 1 million token context and hybrid attention (OpenRouter, accessed 2026-08-02), is the realistic self-hosting tier for most teams. Reserve frontier capability for the specific workloads that fail without it, and buy it rather than hold it. Source confidence: Analytical.
Sources
- DeepSeek V4 Flash release date 2026-04-24, 284 billion total and 13 billion active parameters, 1 million token context, hybrid attention, high and maximum reasoning effort levels, and pricing of $0.0868 input and $0.1736 output per million tokens at a 38 percent promotional discount: OpenRouter model page (accessed 2026-08-02). Primary.
- Joint UK AI Security Institute and United States Center for AI Standards and Innovation preliminary assessment of Kimi K3 cyber capabilities, including "The Last Ones" step counts, ExploitBench scores against GLM-5.2, arbitrary code execution results, and the safeguards finding: UK AI Security Institute (2026-07-23). Primary.
- Kimi K3 weight footprint, four-bit and 16-bit storage requirements, accelerator implication, and the custom Kimi K3 License: techi.com (2026-07) and qz.com (2026-07). Mixed: named public reporting, not a primary institutional record.
- Kimi K3 release on 2026-07-16, open weights shipped 2026-07-26 to 27 on Hugging Face in 96 shards, roughly 2.8 trillion total and 104 billion active parameters (16 of 896 experts), 1 million token context, native vision: Moonshot model card, corroborated by the joint institute assessment and by independent release coverage. Mixed: the originating page was not opened on 2026-08-02; the figures are independently corroborated.
Unverified as of 2026-08-02: the Kimi K3 technical paper and independent cost, speed, and latency profiles were not opened this run. No cost, speed, or latency figure for Kimi K3 is asserted in this article.
Methodology
This article follows the Bottleneck Map method. The bottleneck is assigned to Layer 3, Compute & AI Infrastructure, because the primary constraint is the accelerator and memory footprint required to serve frontier open weights, not model availability, licensing enforcement, or export policy. A license constraint is present and noted, but it is downstream of the deployment economics that decide who can host at all.
Source confidence is stated per claim: Primary where a named, publicly verifiable institutional source is cited inline; Mixed where public reporting is combined with editorial interpretation; Analytical where the claim reflects Stack & State ecosystem observation. Analytical classifications are hypotheses for navigation, not verified factual findings, and should not be used as legal, investment, procurement, or compliance advice. All evaluation scores are specific to the evaluation design that produced them and are reported as such.
Research cutoff and access date: 2026-08-02. Corrections: /connect/.
Stack & State is an editorial and ecosystem-intelligence publication. Nothing here is legal, investment, procurement, or compliance advice. Program details change; verify requirements with primary sources and qualified advisors.
Verified sources
Last verified