Kimi K2.5
Kimi K2.5 is a multimodal model with a 262,144-token context window. It is closed to new API users and the platform sunsets it on August 31, 2026.
Last updated
kimi-k2.5Key facts
At a glance
- Context window
- 262,144
- Parameters
- 1T
- Released
- January 1, 2026
- Output / 1M
- $3.00
Pricing
How much does Kimi K2.5 cost?
Kimi K2.5 costs $0.60 per million input tokens on a cache miss, $0.10 per million on a cache hit, and $3.00 per million output tokens. Prices exclude tax.
| Token type | Price per 1M tokens |
|---|---|
| Input — cache hit | $0.10 |
| Input — cache miss | $0.60 |
| Output | $3.00 |
What a real request costs
A 100,000-token prompt with a 5,000-token response costs $0.075 cold. Repeat that same prompt prefix so 80% of the input hits the cache and it drops to $0.035 — a 53% saving.
This is why headline per-token prices mislead on agent workloads, where the system prompt and tool definitions repeat on every call.
Capabilities
What can Kimi K2.5 do?
- Text inputdocumented
- Image inputdocumented
- Video inputdocumented
- Thinking modedocumented
- Tool callingdocumented
- JSON modedocumented
- Structured outputnot documented
- Partial modedocumented
- Context cachingdocumented
- Web searchdocumented
- Open weightsdocumented
Notes
What the spec sheet leaves out
- Not available to accounts registered after the Kimi K3 launch.
- Full platform sunset is scheduled for August 31, 2026 — existing users should migrate.
- Still the cheapest Kimi model per token while it remains available.
- Moonshot has published the month but not an exact release day; January 2026 is what we can confirm.
Analysis
Choosing Kimi K2.5
Can you still use Kimi K2.5?
Only if your account already has access. Moonshot AI closed Kimi K2.5 to accounts registered after the Kimi K3 launch, and sunsets the model entirely on August 31, 2026. For most people reading this, it is not a purchasing option — it is a migration problem or a historical note.
That makes this page unusual. Every price below is real and currently in force, and almost none of it should influence a decision you make today. If you are choosing a model, start at Kimi K2.6 or Kimi K3 instead.
If you do have an existing integration, the rest of this page is for you, and the useful part is the timeline.
What should you migrate to?
Kimi K2.6 for almost everyone. It matches K2.5 on the things that usually break a migration — same 262,144-token context window, same text, image and video input, web search still documented — and it has no scheduled retirement.
| Kimi K2.5 | Kimi K2.6 | Kimi K2.7 Code | |
|---|---|---|---|
| Context window | 262,144 | 262,144 | 262,144 |
| Input, cache hit | $0.10 | $0.16 | $0.19 |
| Input, cache miss | $0.60 | $0.95 | $0.95 |
| Output | $3.00 | $4.00 | $4.00 |
| Web search | Documented | Documented | Not documented |
| Status | Sunsets August 31, 2026 | Current | Current |
Per million tokens.
Choose K2.7 Code instead if the workload is mostly code — it is priced identically to K2.6 on cache-miss input and output, and it is tuned for instruction-following in long contexts. The one thing it gives up is web search.
Go to K3 only if you need more than 262,144 tokens of context or structured output. Neither is a reason most K2.5 workloads will hit, and K3 costs considerably more per token.
Our migration guide walks through the actual code changes, which in most cases amount to changing one model identifier.
Your bill will go up. How much?
Expect an increase on output tokens, because K2.5 is the cheapest model Moonshot still serves. Output goes from $3.00 to $4.00 per million on the move to K2.6, and cache-miss input from $0.60 to $0.95.
We would rather state that plainly than describe the migration as a free upgrade. It is not. The price difference is real and it is the reason people delay the move.
What softens it is cache behaviour. Both models discount cached input by roughly the same proportion, so a workload with a stable repeated prefix sees a smaller effective increase than the headline rates suggest. If you have not audited your cache hit rate, the migration is a good moment to — the cost calculator models both models side by side with a cache-hit input.
Why is the deadline the part that matters?
Because the alternative to migrating on your schedule is migrating on Moonshot’s, in production, after requests start failing. Moonshot has not published what happens to K2.5 calls after August 31, 2026, so the safe assumption is that they stop working rather than quietly routing somewhere else.
There is precedent worth taking seriously. The whole kimi-k2-* series was discontinued
on 25 May 2026 and those endpoints no longer respond — no fallback, no grace period we are
aware of. Kimi K2 is the page for that, and it is the second Kimi
retirement inside four months. We wrote up what this shutdown means and what the pattern
suggests in Moonshot is retiring Kimi K2.5 on 31 August 2026.
That cadence is the real lesson here, and it is why this site puts a verification date next to every number. Moonshot ships and retires models faster than most published comparisons get updated. A price you find quoted elsewhere for a Kimi model may be describing something that no longer exists.
What does the migration actually involve?
For most integrations, changing one string. The Kimi API is Chat Completions-shaped across every model, so the request and response formats do not change between K2.5 and K2.6 — which is why this migration is far less work than the deadline makes it feel.
The checklist we would work through:
- Find every hardcoded model identifier.
kimi-k2.5may appear in more than one place — application code, background jobs, evaluation scripts, notebooks. This is the step people miss, and the leftovers fail silently until the shutdown date. - Move it into configuration. If you are touching every call site anyway, make the next migration a deploy rather than a refactor. There will be a next migration.
- Check capability assumptions. K2.6 documents the same text, image and video input and keeps web search, so most code carries over untouched. Confirm rather than assume if you rely on anything unusual.
- Re-run your evaluations. Different model, potentially different outputs. If you have no evaluation set, a sample of real prompts compared side by side is better than nothing.
- Re-estimate cost before you switch, not after. The rates go up. Model it in the cost calculator with your real cache hit rate so the first invoice is not a surprise.
- Audit your cached prefix while you are in there. Cache-hit pricing is where the increase gets absorbed, and a prefix with a timestamp interpolated into it is not getting cached at all.
The migration guide covers the code-level detail.
What if you cannot migrate in time?
The weights are downloadable, so self-hosting is technically an option — K2.5 is published as open weights, and the API shutdown does not remove them. That is a genuine escape hatch for a workload that cannot move.
It is also a serious undertaking, and it is not the same as a hosted API. You take on serving infrastructure, throughput, availability and the licence terms attached to the weight release. Read that licence before assuming commercial deployment is permitted. For most teams, changing one model identifier is dramatically cheaper than standing up inference.
FAQ
Kimi K2.5 questions
What is Kimi K2.5?
Kimi K2.5 is a multimodal model with a 262,144-token context window. It is closed to new API users and the platform sunsets it on August 31, 2026.
What is the Kimi K2.5 context window?
Kimi K2.5 has a 262,144-token context window. That total covers input and output combined, so a long prompt reduces the room left for the response.
How much does Kimi K2.5 cost?
Kimi K2.5 costs $0.60 per million input tokens on a cache miss, $0.10 per million on a cache hit, and $3.00 per million output tokens. Prices exclude tax.
When does Kimi K2.5 shut down?
Moonshot AI sunsets Kimi K2.5 on August 31, 2026. It is already closed to newly registered accounts, so plan a migration to Kimi K3 or Kimi K2.7 Code before that date.
Is Kimi K2.5 open source?
Moonshot AI published downloadable open weights for Kimi K2.5, so you can self-host it. That is not the same as an OSI-approved open-source licence — check the licence attached to the weight release before deploying commercially.
Can I still sign up for Kimi K2.5?
No. Moonshot AI closed Kimi K2.5 to accounts registered after the Kimi K3 launch. If your account is newer than that, the model is not available to you at any price, and it is scheduled for full shutdown on 31 August 2026 regardless.
What should I migrate to from Kimi K2.5?
Kimi K2.6 is the closest match — same 262,144-token context window, same multimodal input, still has web search, and no scheduled retirement. Move to Kimi K2.7 Code instead if the workload is primarily coding, since it costs the same on output tokens.
What happens to my requests after the Kimi K2.5 sunset date?
Moonshot has not published the exact failure behaviour after 31 August 2026. When Kimi K2 was discontinued in May 2026 those endpoints stopped responding. Plan for requests to fail rather than to silently route elsewhere, and migrate before the date.
Sources
Where these figures come from
Every specification and price on this page is transcribed from Moonshot AI’s own documentation and re-checked on the date shown. Nothing is estimated. Our editorial policy sets out how we source figures and what we do when something cannot be verified.