Moonshot AI released Kimi K3 on 16 July 2026 and published the weights eleven days later under a modified MIT licence. It is a 2.8 trillion parameter mixture of experts model with a 1 million token context, a new attention scheme Moonshot calls Kimi Delta Attention, and native vision. Moonshot positions it second only to Claude Fable 5 and GPT-5.6 Sol on GDPval AA v2, so the comparison in this paper is the one Moonshot chose for itself.
The received wisdom is that Chinese models are the cheap option and American models are the expensive one. We priced it properly, and that is no longer true at the top of the range. Kimi K3 costs more per token than Claude Sonnet 5. Moonshot moved it there deliberately, and understanding why explains more about where the industry is going than any benchmark table does.
What follows is the cost audit, the self hosting maths, and what the pricing reveals about two labs running opposite strategies for reasons that have very little to do with philosophy.
1. The sticker price
All rates are per million tokens, taken from the providers rather than from aggregators, and current at the start of September 2026.
| Model | Input | Output | Cached input | Context |
|---|---|---|---|---|
| Claude Fable 5 | $10.00 | $50.00 | about $1.00 | 1M |
| Claude Opus 5 | $5.00 | $25.00 | about $0.50 | 1M |
| Claude Sonnet 5 | $2.00 | $10.00 | about $0.20 | 1M |
| Kimi K3 | $3.00 | $15.00 | $0.30 | 1M |
| Kimi K3, cheapest third party host | $2.55 | $12.75 | $0.256 | 1M |
| Kimi K2.6, the previous flagship | $0.95 | $4.00 | n/a | 256K |
Take a mid sized production workload: 50 million input tokens and 5 million output tokens a month, which is a busy internal assistant or a modest agent fleet.
| Model | Monthly | Against K3 |
|---|---|---|
| Claude Fable 5 | $750.00 | 3.33x |
| Claude Opus 5 | $375.00 | 1.67x |
| Kimi K3 | $225.00 | 1.00x |
| Claude Sonnet 5 | $150.00 | 0.67x |
| Kimi K2.6 | $67.50 | 0.30x |
| GLM 5.2 | $38.90 | 0.17x |
| Nemotron 3 Ultra | $34.20 | 0.15x |
| DeepSeek V4 | $26.35 | 0.12x |
| MiniMax M3 | $10.95 | 0.05x |
| DeepSeek V4 Flash | $3.91 | 0.02x |
Fable 5 to K3 is 3.33 times. Fable 5 to DeepSeek V4 Flash is 192 times. Those two numbers are describing completely different markets, and treating them as one story is how procurement decisions go wrong.
2. What the sticker price hides
Three corrections, in order of how much money they move.
Caching does not change the ratio, it changes the bill. Both Anthropic and Moonshot price cached input at roughly a tenth of the standard rate. Run the same workload at a 90% cache hit rate and Fable 5 falls to $345 while K3 falls to $103.50. Still 3.33 times, and two thirds cheaper on both sides. Prompt caching is the single largest lever on either bill and it costs nothing but prefix discipline: stable content first, volatile content last, and a check that cache reads are actually non zero rather than assumed.
Batch halves the gap. Anthropic's Batch API runs at 50%, so Fable 5 batched is $375 against K3 at $225. The gap goes from 3.33 times to 1.67 times. Moonshot has not published batch pricing for K3 at all, though its older models get 60% of standard. For anything that does not need an answer this second, evaluation runs, overnight enrichment, document processing, 1.67 times is the honest comparison, not 3.33.
Effort moves more than the vendor does. Fable 5 reasons on every request and those thinking tokens bill as output, but the depth is a parameter, not a fixed property. Dropping a route from high to medium effort can save more than switching providers, and lower effort on a current model often beats high effort on the previous generation. The metric that matters is cost per completed task. A model that is three times cheaper per token and needs three times as many turns has saved nobody anything, and the only way to know which you have is to measure turns to completion on your own work.
3. The finding worth sitting with
Kimi K3 is more expensive than Claude Sonnet 5. Three dollars against two on input, fifteen against ten on output.
It is also a 3.2 times price rise over Moonshot's own previous flagship. Kimi K2.6 shipped at $0.95 and $4.00. K3 shipped at Sonnet money. The leading Chinese lab looked at the top of its range and chose to stop competing on price there.
That is the opposite of what the last two years trained everyone to expect, and it points at a market that has split in two rather than converged:
- A volume tier, where DeepSeek V4 Flash, MiniMax M3 and GLM 5.2 are driving the marginal price of a competent token towards zero. This is where the price war actually is, and where a small company's economics genuinely change.
- A frontier tier, where Kimi K3 now sits alongside the Western models and is priced like them, because at the frontier the constraint is capability and the buyer is not shopping on unit price.
K3 is not in the price war. It left.
4. Open weights, priced honestly
The weights are free to download. For almost every business reading this, they are also unusable, and it is worth being precise about why rather than repeating that open weights mean cheaper inference.
At 2.8 trillion parameters in MXFP4, the weights alone are about 1.4 terabytes. That is roughly ten high memory GPUs before a single byte of KV cache, and a 1 million token context wants a great deal of KV cache. One serving replica realistically means 16 to 24 GPUs.
At typical rental rates of two to three dollars per GPU hour, that is $23,000 to $53,000 a month for one replica that has to stay warm. Against Moonshot's own API at a blended rate of about $4.09 per million tokens, break even lands somewhere between 5.7 and 12.8 billion tokens a month. Unless you are a neocloud or a sovereign deployment, you will not get there.
So what are open weights worth? Not a lower bill. They are worth auditability, portability and the absence of a vendor who can deprecate your model. Those are real and in some sectors decisive, but they are governance properties, not cost properties, and they should be argued for on those terms. The models where open weights genuinely rewrite a small company's cost base are the small ones, not the 2.8 trillion parameter flagship. We looked at the same economics from the hardware side in the GPU query engine paper: utilisation, not list price, decides whether owning compute beats renting it.
One more thing to check before you download anything. Open weight is four different legal positions right now. DeepSeek V4 and GLM 5.2 ship under MIT. Kimi K3 is modified MIT. MiniMax M3 uses a restricted community licence. NVIDIA's Nemotron uses OpenMDW. Read the file. And if you modify an open weight model and put it on the EU market under your own name, look carefully at where that leaves you under the AI Act, because the roles shift with the branding, as we set out in the EU AI Act guide.
5. Why the two sides are running opposite strategies
This is not a values difference. It is what each side's constraints make rational.
The binding constraint is compute. US hyperscaler capital expenditure was at least $350 billion in 2025, against under $40 billion for China's major cloud providers. On 31 May 2026 the US Bureau of Industry and Security extended licensing to any China parented buyer regardless of where the subsidiary sits, closing the Singapore and Malaysia neocloud channel for new purchases. Efficiency stopped being a preference and became the only available strategy, which is why the architectural signature of the Chinese frontier is sparsity and low precision: mixture of experts, new attention schemes like Kimi Delta Attention, and a checkpoint shipped natively in MXFP4 rather than quantised down afterwards.
Open weights are a distribution strategy, not a gift. You cannot win a closed API race from a ninefold capital disadvantage. What you can do is commoditise the layer your competitor monetises, and compete somewhere they are not. Open weights put a model into every neocloud, every enterprise VPC and every sovereign deployment without ever winning on API price. If the world's agent scaffolding ends up written against your tool calling format, you have taken the platform layer while the closed labs were defending the capability lead.
It is working, measurably. Stanford's 2026 AI Index put the gap between the leading US and Chinese models at 2.7% as of March 2026. OpenRouter's routing data shows Chinese open weight models moving from a negligible share to a majority of all tokens processed between late 2024 and mid 2026. And the open weight lag behind the closed frontier has held steady at three to six months for more than eighteen months, which is the number that should worry a closed lab: the frontier is not pulling away.
The tell is NVIDIA. The strongest American open weight model is Nemotron, from the company that sells the hardware. Its incentive is structurally identical to Moonshot's: commoditise the model to sell the complement. Chips in one case, a national ecosystem in the other, same play. Whenever a company gives away something valuable, the question is what it sells that the giveaway makes more valuable.
6. What the American labs are competing on instead
Look at what Anthropic ships against a 3.33 times price gap and it is not a price cut.
It is cached reads at a tenth of standard, batch at half, effort control, task budgets, and a managed agent runtime. Every one of those attacks cost per completed task rather than cost per token. That is the coherent answer to being undercut: compete on the total cost of finishing the job, where a model that needs fewer turns, retries and human corrections wins even at three times the unit price.
Whether it holds is an empirical question and it will be settled workload by workload, not by argument. Which is exactly why the second flow above starts by measuring the cheap tier against the quality bar rather than by picking a model.
7. What we would actually do
Four steps, in this order, and the first two are free.
- Fix caching before comparing anything. If your cache read count is zero, you are comparing two badly run bills. A timestamp in a system prompt or an unsorted JSON blob is usually the culprit.
- Move everything that can wait to batch. Half price on the Claude side, and it changes which comparison you are even having.
- Start at the volume tier, not the frontier. Run DeepSeek V4 Flash or GLM 5.2 against your quality bar first. If they clear it, the frontier tier debate is moot and you have saved an order of magnitude. If they do not, you now know what you are paying the frontier premium to fix.
- Score cost per completed task, never per token. Turns to completion, retry rate and human correction time all belong in the number. This is the only measurement that settles the argument, and almost nobody runs it.
The broader pattern is the one we keep meeting. As we argued in the TimesFM 3 note, a 330 million parameter model beating one eight times its size, the capable model is very often not the biggest or the most expensive one, and the constraint that decides adoption is rarely the one in the announcement.
If you want this run properly against your own workload rather than against a benchmark suite, that is a week of work and it usually pays for itself immediately. Talk to us, or see how we build systems for businesses.
8. Sources
- Kimi K3 on OpenRouter, for third party hosted pricing and context limits.
- Moonshot Kimi API pricing, for first party input, output and cached input rates across the Kimi range.
- Moonshot ships Kimi K3 at 2.8T parameters, for the release dates, licence and positioning.
- OpenRouter, the open weight models that matter, June 2026, for comparative pricing and licence terms across DeepSeek, GLM, MiniMax and Nemotron.
- CNBC, China's open weight model lead.
- US China Economic and Security Review Commission, Two Loops: how China's open AI strategy reinforces its industrial dominance.
- Palladium, American AI may not survive Chinese open source, for the capital expenditure and export control position.
- Anthropic model pricing is taken from the current Claude API reference.
Prices and licences in this market move monthly, and everything here is current at 1 September 2026. The self hosting figures are estimates built on stated assumptions about GPU count and rental rate, and they are meant to show where break even sits, not to substitute for a quote. Confirm current rates before making a procurement decision.
This note sits in our Research track, alongside the foundation models and operations threads. Applied work lives in Operations, and the full library is at Research.