Skip to content

Kimi K2.7 Code

Kimi K2.7 Code is Moonshot AI’s coding-focused model, with a 262,144-token context window and higher instruction-following reliability in long contexts.

Last updated

Available nowkimi-k2.7-code

Key facts

At a glance

Context window
262,144
Parameters
1T
Released
Not published
Output / 1M
$4.00

Pricing

How much does Kimi K2.7 Code cost?

Kimi K2.7 Code costs $0.95 per million input tokens on a cache miss, $0.19 per million on a cache hit, and $4.00 per million output tokens. Prices exclude tax.

Source: Moonshot AI platform documentation, verified August 6, 2026.
Token typePrice per 1M tokens
Input — cache hit$0.19
Input — cache miss$0.95
Output$4.00

What a real request costs

A 100,000-token prompt with a 5,000-token response costs $0.115 cold. Repeat that same prompt prefix so 80% of the input hits the cache and it drops to $0.054 — a 53% saving.

This is why headline per-token prices mislead on agent workloads, where the system prompt and tool definitions repeat on every call.

Capabilities

What can Kimi K2.7 Code do?

  • Text inputdocumented
  • Image inputdocumented
  • Video inputdocumented
  • Thinking modedocumented
  • Tool callingdocumented
  • JSON modedocumented
  • Structured outputnot documented
  • Partial modedocumented
  • Context cachingdocumented
  • Web searchnot documented
  • Open weightsdocumented

Notes

What the spec sheet leaves out

  • Moonshot has not published a precise release date for K2.7 Code.
  • Accepts text, image and video input.
  • Open weights (1T total parameters, 32B active) are published at huggingface.co/moonshotai/Kimi-K2.7-Code.

Analysis

Choosing Kimi K2.7 Code

When should you use Kimi K2.7 Code instead of Kimi K2.6?

For anything code-shaped, use K2.7 Code — and note that it costs the same. Moonshot prices both at $0.95 per million cache-miss input tokens and $4.00 per million output tokens. The coding model is not a premium tier. It is the same price, tuned for a narrower job.

That makes the usual “is the specialist worth it” question disappear. There is no premium to justify. The only pricing difference between the two runs the other way: K2.6 has the cheaper cache-hit rate, at $0.16 per million against K2.7 Code’s $0.19.

That gap is small, but on an agent workload where most input is a repeated cached prefix, cache-hit price is the number that dominates the bill. If your traffic is overwhelmingly cached and only incidentally about code, run the numbers before defaulting to the coding model.

Your work Model Reason
Writing, refactoring, reviewing code Kimi K2.7 Code Purpose-built, same cache-miss and output price
Coding agents that must search the web Kimi K2.6 K2.7 Code does not document web search
Responses that must match a strict schema Kimi K3 Structured output is K3-only across the family
Codebases larger than 262,144 tokens Kimi K3 Only model past this context size
Heavy cached-prefix traffic, mixed subject matter Kimi K2.6 Cheaper per cached input token

What does “coding-focused” actually mean here?

Moonshot’s claim is narrower than “better at code”: K2.7 Code is documented as having higher instruction-following reliability in long contexts. That is a claim about staying on task across a large prompt, not about knowing more languages or writing cleverer algorithms.

The distinction matters when you are choosing. Long-context instruction-following is exactly what breaks first in agent loops — the model reads forty files, and by the time it answers it has drifted from the format you asked for, or quietly stopped honouring a constraint you set at the top. That is the failure this model is tuned against.

If your usage is short prompts asking for a single function, you are unlikely to see the difference. If it is a long harness with tool definitions, file contents and a strict output contract, that is where the tuning is aimed.

We have not independently benchmarked this against K2.6, and we will say so rather than repeat the claim as though we had verified it.

Why is there no web search on Kimi K2.7 Code?

Moonshot does not document a built-in web_search capability for this model, although Kimi K2.6 and Kimi K3 both have one. No reason has been published, so we will not invent one. What matters is the consequence: a coding agent on K2.7 Code cannot look things up on its own.

The capability matrix above treats an undocumented feature as “not documented” rather than “confirmed absent” — Moonshot may simply not have written it up. But you cannot ship against a capability a vendor has not committed to in writing.

The practical workaround is that tool calling is documented, so you can wire in your own search tool and pass results back. That is more work than a built-in, and it puts the retrieval quality in your hands rather than Moonshot’s. Whether that is a downgrade or an upgrade depends on how much you trust your own search stack.

Worth noting that K3’s built-in web_search currently comes with Moonshot’s own warning that it is mid-update and not recommended for reliance — so “K2.7 Code lacks search” is less of a gap today than the capability matrix alone suggests.

Should you pay double for HighSpeed?

Only when latency is the binding constraint on your product, because Kimi K2.7 Code HighSpeed is exactly twice the price on every tier for identical weights. Same model, same outputs, faster serving. Nothing about quality changes.

The honest test is whether a user is sitting there watching the tokens arrive. Interactive completion in an editor, a chat surface where someone is waiting — those are latency-bound, and the markup buys something real. Batch jobs, nightly runs, background agents and CI tasks are not, and there the second copy of your bill buys nothing at all.

Kimi K2.7 Code vs HighSpeed works the trade through properly, including when the doubling is defensible.

How do you keep a coding agent inside the context window?

By budgeting the window deliberately rather than filling it and hoping. K2.7 Code holds 262,144 tokens covering input and output combined, and a coding agent will consume that faster than almost any other workload — file contents, tool definitions, diffs and conversation history all accumulate.

The failure is rarely a hard error. What usually happens first is quality degradation as the useful signal gets diluted, followed by the model losing the instruction you gave it forty files ago. The long-context instruction-following this model is tuned for delays that, but it does not remove it.

Three things that help, roughly in order of effect:

  • Send file excerpts, not whole files. Most tasks need a function and its imports, not four hundred lines of unrelated code.
  • Summarise conversation history rather than replaying it. An agent that appends every previous turn verbatim will fill the window with its own transcript.
  • Keep tool definitions stable and near the front. They repeat on every call, so they belong in the cached prefix — which makes them nearly free rather than a recurring cost.

If a task genuinely needs a whole repository in context, that is the case for Kimi K3 and its 1,048,576-token window. Paying K3 prices to avoid doing retrieval properly is usually the more expensive answer, but for whole-codebase reasoning there is no substitute in this family.

What do the open weights actually give you?

Moonshot publishes downloadable weights for this model — 1T total parameters with 32B active — which means self-hosting is possible if the economics or the data-residency rules favour it.

Two caveats we would rather state than let you discover. Downloadable is not the same as open source: read the licence attached to the weight release before deploying commercially or redistributing anything derived from it. And HighSpeed is a serving tier rather than a separate weight release, so there is no separate repository to fetch — if you self-host, your throughput is a function of your own hardware, not of which API tier you would have bought.

FAQ

Kimi K2.7 Code questions

What is Kimi K2.7 Code?

Kimi K2.7 Code is Moonshot AI’s coding-focused model, with a 262,144-token context window and higher instruction-following reliability in long contexts.

What is the Kimi K2.7 Code context window?

Kimi K2.7 Code has a 262,144-token context window. That total covers input and output combined, so a long prompt reduces the room left for the response.

How much does Kimi K2.7 Code cost?

Kimi K2.7 Code costs $0.95 per million input tokens on a cache miss, $0.19 per million on a cache hit, and $4.00 per million output tokens. Prices exclude tax.

Is Kimi K2.7 Code open source?

Moonshot AI published downloadable open weights for Kimi K2.7 Code, so you can self-host it. That is not the same as an OSI-approved open-source licence — check the licence attached to the weight release before deploying commercially.

Is Kimi K2.7 Code better than Kimi K2.6 for coding?

Yes, and it costs the same on cache-miss input and output tokens. Kimi K2.7 Code is purpose-built for coding with better instruction-following in long contexts, and Moonshot prices it identically to K2.6 on those two tiers. For coding work there is no cost argument for K2.6.

Does Kimi K2.7 Code support web search?

No. Moonshot does not document web search for Kimi K2.7 Code, though Kimi K2.6 and Kimi K3 both have it. If your coding agent needs to look things up, you either use a different model or wire in your own search tool through tool calling.

What is the difference between Kimi K2.7 Code and K2.7 Code HighSpeed?

Serving speed and price only. They are the same weights. HighSpeed runs at roughly 180 tokens per second and costs exactly double on every pricing tier. Nothing about output quality changes.

Sources

Where these figures come from

Every specification and price on this page is transcribed from Moonshot AI’s own documentation and re-checked on the date shown. Nothing is estimated. Our editorial policy sets out how we source figures and what we do when something cannot be verified.