Kimi K2.6
Kimi K2.6 is a general-purpose model with a 262,144-token context window, supporting text, image and video input across thinking and non-thinking modes.
Last updated
kimi-k2.6Key facts
At a glance
- Context window
- 262,144
- Parameters
- 1T
- Released
- April 20, 2026
- Output / 1M
- $4.00
Pricing
How much does Kimi K2.6 cost?
Kimi K2.6 costs $0.95 per million input tokens on a cache miss, $0.16 per million on a cache hit, and $4.00 per million output tokens. Prices exclude tax.
| Token type | Price per 1M tokens |
|---|---|
| Input — cache hit | $0.16 |
| Input — cache miss | $0.95 |
| Output | $4.00 |
What a real request costs
A 100,000-token prompt with a 5,000-token response costs $0.115 cold. Repeat that same prompt prefix so 80% of the input hits the cache and it drops to $0.052 — a 55% saving.
This is why headline per-token prices mislead on agent workloads, where the system prompt and tool definitions repeat on every call.
Capabilities
What can Kimi K2.6 do?
- Text inputdocumented
- Image inputdocumented
- Video inputdocumented
- Thinking modedocumented
- Tool callingdocumented
- JSON modedocumented
- Structured outputnot documented
- Partial modedocumented
- Context cachingdocumented
- Web searchdocumented
- Open weightsdocumented
Notes
What the spec sheet leaves out
- The cheapest cache-hit input price in the current lineup at $0.16 per million tokens.
- Same cache-miss and output pricing as Kimi K2.7 Code, so K2.7 Code is the better default for coding work at effectively the same cost.
Analysis
Choosing Kimi K2.6
Where does Kimi K2.6 still win?
On price for general work, and on breadth of input. It has the cheapest cache-hit input rate of any model currently open to new accounts, at $0.16 per million tokens, and it documents text, image and video input — the last of which Kimi K3 does not.
That combination makes K2.6 the sensible default rather than the budget option. The flagship is more capable in two specific ways, and if your workload does not touch either of them you are paying 3.8× more per output token for nothing.
| Kimi K2.6 | Kimi K3 | |
|---|---|---|
| Context window | 262,144 | 1,048,576 |
| Input, cache hit | $0.16 | $0.30 |
| Input, cache miss | $0.95 | $3.00 |
| Output | $4.00 | $15.00 |
| Video input | Documented | Not documented |
| Structured output | Not documented | Documented |
| Non-thinking mode | Available | Always reasons |
Per million tokens. Both support text, image, tool calling, JSON mode, context caching, web search and open weights. Our full head-to-head, Kimi K3 vs Kimi K2.6, reaches a verdict rather than leaving it at “it depends”.
When should you move up to Kimi K3?
Two triggers, and only two: your prompts genuinely exceed 262,144 tokens, or you need responses that conform to a schema you supply. Everything else K2.6 does at a fraction of the cost.
The context trigger is easy to test — count your tokens and see. The context window visualizer translates the limits into pages and lines of code if the raw numbers are hard to picture.
The structured output trigger is more often misunderstood. K2.6 documents JSON mode, which guarantees syntactically valid JSON. It does not guarantee your schema. If you are currently validating K2.6 responses against a schema and retrying on failure, that retry loop has a cost, and at some volume K3’s structured output becomes cheaper than the retries even at the higher token price. Worth measuring rather than assuming in either direction.
Why does Kimi K2.7 Code cost the same as Kimi K2.6?
Because Moonshot prices them identically on cache-miss input and output — $0.95 and $4.00 per million respectively. The coding model is not a premium product. That has a blunt consequence: for code, Kimi K2.7 Code is the better default at no extra cost.
The one place K2.6 keeps an edge is cached input, where it is $0.16 against K2.7 Code’s $0.19 per million. On an agent workload dominated by a repeated cached prefix, that is the tier that matters most.
The other place K2.6 wins is web search. Moonshot documents it for K2.6 and not for K2.7 Code, so a coding agent that needs to look things up on its own is a K2.6 job — or a tool-calling job, wiring in your own search.
How much does the cache discount actually save?
A cached input token on K2.6 costs 83% less than an uncached one. That is a large enough gap that prompt-prefix discipline is worth designing around rather than treating as an optimisation to get to later.
The mechanism rewards byte-identical prefixes. A system prompt and a block of tool definitions that repeat unchanged across calls will land in cache. The same prefix with a timestamp, a request id or a reshuffled tool order interpolated into it will not, and every call pays $0.95 per million instead of $0.16.
This is the single most common way we see agent bills come in higher than the estimate. The fix is usually one line — move the volatile value out of the prefix and into the user message, where it belongs anyway. Kimi API pricing explained covers what you are actually billed for, cache tiers included.
Is Kimi K2.6 the better choice for web search right now?
Possibly, and for a reason that has nothing to do with K2.6 being better at it. Moonshot
documents web_search for K2.6, and also documents it for Kimi K3 —
but notes that K3’s implementation is mid-update and does not currently recommend relying
on it. That is Moonshot’s own caveat, not ours.
So on the one capability where you might expect the flagship to lead, the vendor is telling you to be careful. Kimi K2.7 Code does not document web search at all, which leaves K2.6 as the current model with search documented and no warning attached to it.
We have not tested either implementation, so we are reporting Moonshot’s position rather than a result of our own. Treat this as a reason to evaluate before committing, not as a verdict. If search grounding is load-bearing in your product, test it on your own queries regardless of which model you choose.
When should you turn thinking mode off?
Whenever the task does not need reasoning — which is more often than people assume. K2.6
documents both thinking and non-thinking modes, and that choice is something
Kimi K3 does not offer at all. K3 always reasons; its reasoning_effort
field goes down to low but never to off.
That makes K2.6 the right model for a whole class of work that K3 handles expensively. Classification, extraction, short rewrites, routing decisions, format conversion — none of these benefit from a reasoning pass, and on K3 you pay for one on every single request whether it helped or not.
Reasoning tokens are billed as output tokens. So a workload of ten thousand short classification calls pays for ten thousand reasoning passes it never needed. On a model charging $15.00 per million output tokens rather than K2.6’s $4.00, that compounds quickly.
The pattern worth adopting is to split traffic by task rather than picking one model for the whole application. Route the genuinely hard reasoning to a thinking mode, and let the mechanical work run without one. This is the same argument as the HighSpeed tiering on Kimi K2.7 Code — the expensive option is chosen per request, not per project.
What about Kimi K2.5, which is cheaper still?
It is cheaper, and it is not an option for most readers. Kimi K2.5 undercuts K2.6 at $0.10 cache-hit and $3.00 output per million — but it is closed to accounts registered after the Kimi K3 launch, and Moonshot sunsets the model entirely on August 31, 2026.
If you have an existing K2.5 integration, K2.6 is the natural landing place: same context window, same multimodal input, still has web search, and it is not scheduled for retirement. Our migration guide covers the move.
If you do not already have K2.5 access, the price is academic — you cannot buy it.
FAQ
Kimi K2.6 questions
What is Kimi K2.6?
Kimi K2.6 is a general-purpose model with a 262,144-token context window, supporting text, image and video input across thinking and non-thinking modes.
What is the Kimi K2.6 context window?
Kimi K2.6 has a 262,144-token context window. That total covers input and output combined, so a long prompt reduces the room left for the response.
How much does Kimi K2.6 cost?
Kimi K2.6 costs $0.95 per million input tokens on a cache miss, $0.16 per million on a cache hit, and $4.00 per million output tokens. Prices exclude tax.
Is Kimi K2.6 open source?
Moonshot AI published downloadable open weights for Kimi K2.6, so you can self-host it. That is not the same as an OSI-approved open-source licence — check the licence attached to the weight release before deploying commercially.
Is Kimi K2.6 still worth using now that Kimi K3 exists?
Yes, for most general work. Kimi K2.6 costs substantially less per token than Kimi K3 on every tier and documents video input, which K3 does not. Move up to K3 only when you need a context window beyond 262,144 tokens or structured output.
What is the cheapest Kimi model for cached input?
Kimi K2.6, at $0.16 per million cache-hit input tokens, is the cheapest model in the current lineup. Kimi K2.5 is cheaper still at $0.10, but it is closed to new API accounts and Moonshot sunsets it on 31 August 2026.
Does Kimi K2.6 support structured output?
No. Moonshot does not document structured output for Kimi K2.6. JSON mode is documented, which constrains responses to valid JSON but does not enforce a schema you supply. Kimi K3 is the only model in the family with documented structured output.
Sources
Where these figures come from
Every specification and price on this page is transcribed from Moonshot AI’s own documentation and re-checked on the date shown. Nothing is estimated. Our editorial policy sets out how we source figures and what we do when something cannot be verified.