Kimi K3
Kimi K3 is Moonshot AI’s flagship model, with 2.8 trillion parameters, native visual understanding and a 1,048,576-token context window.
Last updated
kimi-k3Key facts
At a glance
- Context window
- 1,048,576
- Parameters
- 2.8T
- Released
- July 16, 2026
- Output / 1M
- $15.00
Pricing
How much does Kimi K3 cost?
Kimi K3 costs $3.00 per million input tokens on a cache miss, $0.30 per million on a cache hit, and $15.00 per million output tokens. Prices exclude tax.
| Token type | Price per 1M tokens |
|---|---|
| Input — cache hit | $0.30 |
| Input — cache miss | $3.00 |
| Output | $15.00 |
What a real request costs
A 100,000-token prompt with a 5,000-token response costs $0.375 cold. Repeat that same prompt prefix so 80% of the input hits the cache and it drops to $0.159 — a 58% saving.
This is why headline per-token prices mislead on agent workloads, where the system prompt and tool definitions repeat on every call.
Capabilities
What can Kimi K3 do?
- Text inputdocumented
- Image inputdocumented
- Video inputnot documented
- Thinking modedocumented
- Tool callingdocumented
- JSON modedocumented
- Structured outputdocumented
- Partial modedocumented
- Context cachingdocumented
- Web searchdocumented
- Open weightsdocumented
Notes
What the spec sheet leaves out
- Always reasons. Reasoning depth is set with the top-level
reasoning_effortfield:low,highormax. The default ismax. - Adds two API capabilities over earlier models: tool choice constraints (
tool_choice) and dynamically loaded tools. - Moonshot notes that
web_searchis mid-update and does not currently recommend relying on it. - Open weights were published in July 2026.
Analysis
Choosing Kimi K3
When should you actually use Kimi K3?
Reach for Kimi K3 when the job genuinely needs more than 262,144 tokens of context, or when you need structured output. On everything else Kimi K2.6 does the same work for less money. K3 is a specialist that happens to also be the flagship, and those are different things.
The registry figures above make the trade explicit. K3 carries 2.8T parameters against K2.6’s 1T, and a 1,048,576-token window against 262,144 — four times the room. You pay for both. Output tokens cost $15.00 per million versus $4.00, and that gap applies to every response whether or not you used the extra context.
So the question is not “which model is better”. It is whether your workload actually touches the thing you are paying extra for. We work through that decision in detail in Kimi K3 vs Kimi K2.6, including where each one wins.
| Your workload | Model | Why |
|---|---|---|
| Whole-repository or whole-corpus reasoning | Kimi K3 | Only model in the family past 262,144 tokens |
| Responses that must match a strict schema | Kimi K3 | K2.6 does not document structured output |
| Video input | Kimi K2.6 or K2.7 Code | K3 does not document video; both of those do |
| General chat, summarisation, extraction | Kimi K2.6 | Same capability set, lower price |
| Coding | Kimi K2.7 Code | Purpose-built, and cheaper than both |
Where does Kimi K3 lose to Kimi K2.6?
Two places, and one of them is surprising. K3 costs more on every pricing axis, which is expected of a flagship. Less expected: Moonshot does not document video input for K3, while K2.6 does. A multimodal pipeline that reads video has to use the older, cheaper model.
The capability matrix above is built from Moonshot’s own documentation, and a missing entry there means “not documented” rather than “confirmed absent” — Moonshot may simply not have published it. But you cannot build on a capability a vendor has not committed to in writing, so for planning purposes the gap is real.
The pricing gap is the other half. K3’s cache-miss input price is $3.00 per million tokens against K2.6’s $0.95. If your prompts are short and varied — the profile of most chat and extraction work — you are paying flagship rates for context you never fill.
How does the million-token context window behave in practice?
The 1,048,576-token window is a combined budget, not an input allowance. Input and output are drawn from the same total, so a prompt that fills most of the window leaves little room for the model to answer. Plan the response length as part of the budget, not as an afterthought.
This catches people migrating from models where the output limit was quoted separately. Moonshot does not publish a separate maximum output length for K3, so the only ceiling is what the context has left. A 1,000,000-token prompt does not leave room for a long answer.
The practical pattern is to keep a deliberate reserve — decide the largest response you need, subtract it, and treat the remainder as your real input budget.
What does reasoning_effort actually change?
Kimi K3 always reasons. reasoning_effort accepts low, high or max, and defaults to
max — so unless you set it, every request pays for the deepest reasoning mode, including
requests that did not need it. This is the single easiest cost mistake to make on K3.
There is no setting that disables reasoning. If your workload is genuinely simple — classification, extraction, short rewrites — that is an argument for using K2.6 rather than for tuning K3 down, because K2.6 offers a non-thinking mode and K3 does not.
Two API capabilities arrived with K3 that earlier models do not have: tool choice
constraints via tool_choice, and dynamically loaded tools. If you are building agents,
those are the reasons to be on K3 that have nothing to do with context size.
Why does cache hit rate matter more on Kimi K3 than anywhere else?
Because K3 has the steepest cache discount in the family. A cached input token costs 90% less than an uncached one, against 83% on Kimi K2.6. The same repeated-prefix discipline saves you proportionally more here than on any other Kimi model.
That has a direct design consequence for agent workloads, where a system prompt and a block of tool definitions repeat on every single call. Keep that prefix byte-identical across requests and it lands in cache. Interpolate a timestamp, a request id or a shuffled tool order into it and every call pays the full $3.00 rate.
The worked example above the fold shows the effect on a single realistic request. Run your own numbers in the API cost calculator — the cache hit rate input is the one that moves the total most — and Kimi API pricing explained covers how the cache tiers are billed.
Are the open weights actually useful to you?
Probably not, and it is worth being honest about why. Moonshot published downloadable weights for K3 in July 2026, which sounds like an escape hatch from API pricing. But this is a 2.8T-parameter model, and serving it yourself is a datacentre problem rather than a server problem.
Compare that with Kimi K2.7 Code, published at 1T total with 32B active. Sparse activation is what makes a model practical to self-host, and the smaller members of the family are far more realistic targets. If your reason for wanting weights is cost control, the cheaper hosted models will beat self-hosted K3 for almost any volume you are likely to have.
The reasons that do hold up are data residency, an air-gapped deployment, or reproducing a result exactly. Those are real, and for them the weights matter regardless of the arithmetic.
One caveat that applies to every open-weights release in this family: downloadable is not open source. The weights ship under a licence attached to the release, and that licence governs commercial deployment and redistribution. We describe every Kimi model as open-weights rather than open-source throughout this site for exactly that reason — the distinction is legal, not pedantic.
What should you watch out for before deploying Kimi K3?
Two things Moonshot has flagged itself. First, web_search is mid-update and Moonshot
does not currently recommend relying on it — so if search grounding is load-bearing in
your product, do not build it on K3’s built-in tool yet.
Second, K3’s weights are downloadable, but downloadable is not the same as open source. Moonshot published open weights in July 2026 under a licence attached to the weight release. Read that licence before deploying commercially or assuming you can redistribute anything derived from it. We describe the model as open-weights rather than open-source throughout this site for exactly that reason.
FAQ
Kimi K3 questions
What is Kimi K3?
Kimi K3 is Moonshot AI’s flagship model, with 2.8 trillion parameters, native visual understanding and a 1,048,576-token context window.
What is the Kimi K3 context window?
Kimi K3 has a 1,048,576-token context window. That total covers input and output combined, so a long prompt reduces the room left for the response.
How much does Kimi K3 cost?
Kimi K3 costs $3.00 per million input tokens on a cache miss, $0.30 per million on a cache hit, and $15.00 per million output tokens. Prices exclude tax.
Is Kimi K3 open source?
Moonshot AI published downloadable open weights for Kimi K3, so you can self-host it. That is not the same as an OSI-approved open-source licence — check the licence attached to the weight release before deploying commercially.
Should I use Kimi K3 or Kimi K2.6?
Use Kimi K2.6 unless you need a context window larger than 262,144 tokens or structured output. K2.6 costs less per token on every axis and documents video input, which Kimi K3 does not. K3 earns its price on long-context work, not on general chat.
Does Kimi K3 support video input?
Moonshot AI does not document video input for Kimi K3. Image input is documented. Kimi K2.6 does document video input, so a pipeline that needs video should use K2.6 rather than the newer flagship.
Can you turn off reasoning in Kimi K3?
No. Kimi K3 always reasons. The reasoning_effort field accepts low, high or max and defaults to max, so you can reduce depth but not disable it. Budget for reasoning tokens on every request, including short ones.
Sources
Where these figures come from
Every specification and price on this page is transcribed from Moonshot AI’s own documentation and re-checked on the date shown. Nothing is estimated. Our editorial policy sets out how we source figures and what we do when something cannot be verified.