Skip to content

Kimi K2.7 Code HighSpeed

Kimi K2.7 Code HighSpeed is the same model as Kimi K2.7 Code served at roughly 180 tokens per second, at double the price.

Last updated

Available nowkimi-k2.7-code-highspeed

Key facts

At a glance

Context window
262,144
Parameters
1T
Released
Not published
Output / 1M
$8.00

Pricing

How much does Kimi K2.7 Code HighSpeed cost?

Kimi K2.7 Code HighSpeed costs $1.90 per million input tokens on a cache miss, $0.38 per million on a cache hit, and $8.00 per million output tokens. Prices exclude tax.

Source: Moonshot AI platform documentation, verified August 6, 2026.
Token typePrice per 1M tokens
Input — cache hit$0.38
Input — cache miss$1.90
Output$8.00

What a real request costs

A 100,000-token prompt with a 5,000-token response costs $0.230 cold. Repeat that same prompt prefix so 80% of the input hits the cache and it drops to $0.108 — a 53% saving.

This is why headline per-token prices mislead on agent workloads, where the system prompt and tool definitions repeat on every call.

Capabilities

What can Kimi K2.7 Code HighSpeed do?

  • Text inputdocumented
  • Image inputdocumented
  • Video inputdocumented
  • Thinking modedocumented
  • Tool callingdocumented
  • JSON modedocumented
  • Structured outputnot documented
  • Partial modedocumented
  • Context cachingdocumented
  • Web searchnot documented
  • Open weightsdocumented

Notes

What the spec sheet leaves out

  • Identical weights to Kimi K2.7 Code — the only difference is serving speed and price.
  • Those weights (1T total parameters, 32B active) are published at huggingface.co/moonshotai/Kimi-K2.7-Code. HighSpeed is a serving tier, not a separate weight release, so there is no HighSpeed-specific repository.
  • Moonshot quotes roughly 180 tokens/second, reaching up to 260 tokens/second in short-context scenarios.
  • Costs exactly 2× Kimi K2.7 Code on every pricing tier. Worth it only when latency is the binding constraint.

Analysis

Choosing Kimi K2.7 Code HighSpeed

What is actually different about HighSpeed?

Serving speed and price. Nothing else. Moonshot publishes the same weights behind both tiers — 1T total parameters, 32B active — so Kimi K2.7 Code and HighSpeed produce the same quality of output from the same prompt. You are buying throughput, not capability.

The capability matrices on the two pages are identical because the underlying model is identical. Same 262,144-token context window, same multimodal input, same tool calling, same absence of web search.

This is unusually clean as a purchasing decision, because there is exactly one variable. Most model choices trade capability against price and force a judgement call. This one does not.

Is HighSpeed worth exactly double?

It costs precisely 2× on every pricing tier — cache-hit input, cache-miss input and output alike. So the question reduces to whether the latency is worth a 100% markup on your entire bill, with no partial credit anywhere.

Tier Kimi K2.7 Code HighSpeed
Input, cache hit $0.19 $0.38
Input, cache miss $0.95 $1.90
Output $4.00 $8.00

Per million tokens. The doubling is exact at every tier, which is itself informative: Moonshot is not pricing a better model, it is pricing a queue position. Our K2.7 Code vs HighSpeed comparison takes the decision apart in more detail.

Moonshot quotes roughly 180 tokens per second for HighSpeed, reaching up to 260 in short-context scenarios. It does not publish a throughput figure for standard K2.7 Code, so we cannot tell you the multiplier — and we are not going to estimate one. If you need to know the real difference for your workload, measure both on your own prompts.

When is latency actually the binding constraint?

When a human is watching tokens arrive. That is the whole test, and it is more restrictive than it first sounds — most production LLM traffic is not being watched by anyone.

Latency-bound, worth considering:

  • Inline completion in an editor, where a pause breaks flow
  • A chat surface where someone is waiting on the first token
  • Interactive agents a developer is actively supervising
  • Anything with a visible progress indicator a user is staring at

Not latency-bound, and here the markup buys nothing:

  • CI pipelines and pre-merge checks
  • Nightly or scheduled batch runs
  • Background agents whose results are read later
  • Bulk generation, migration and backfill jobs
  • Anything already sitting behind a queue

The mistake we would expect people to make is putting a whole application on HighSpeed because part of it is interactive. The tier is chosen per request, not per project. Route the interactive path to HighSpeed and everything else to standard, and you pay the markup only where it does something.

What does this cost in practice?

Because the multiplier is flat, the arithmetic is unusually simple: whatever standard K2.7 Code would cost you, HighSpeed costs twice. There is no volume at which the gap narrows and no cache behaviour that softens it — the cache discount is the same 80% on both tiers, so caching saves you the same proportion of a bill that is twice as large.

Put concretely, a workload spending $500 a month on K2.7 Code spends $1,000 on HighSpeed for identical outputs. The question is whether the latency improvement is worth $500 a month to your users. For an editor plugin with paying customers it might be. For a nightly job nobody watches it is not.

Model your own numbers in the API cost calculator, which carries both tiers as separate models.

How do you test whether HighSpeed actually helps you?

Measure both tiers on your own prompts, because Moonshot publishes a throughput figure for HighSpeed and none for standard K2.7 Code — so the size of the improvement is genuinely unknown until you check it.

A test that produces a usable answer:

  1. Take real prompts, not synthetic ones. Throughput varies with context length, and Moonshot’s own figures say so — roughly 180 tokens per second generally, up to 260 in short-context scenarios. A short benchmark prompt will flatter HighSpeed relative to your actual traffic.
  2. Measure time to first token separately from tokens per second. For interactive use these feel completely different. A user notices the wait before anything appears far more than the speed of the stream once it starts.
  3. Run enough samples to see variance. A handful of calls will not separate the tiers from ordinary fluctuation in load.
  4. Test at your real concurrency. Serving behaviour under parallel load is not the same as one request at a time, and agent workloads are rarely serial.
  5. Convert the result into money. You now have a latency delta and a known 2× cost multiplier. Whether the trade is worth it is a product question, not a technical one.

We have not run this measurement ourselves yet. When we do, the numbers and the method will be published here — and they will be labelled as our measurement rather than as a Moonshot specification, because that distinction is the entire value of doing it.

What should you watch out for?

Self-hosting does not have a HighSpeed option. The published weights are the K2.7 Code weights; HighSpeed is a serving tier on Moonshot’s infrastructure, not a separate release. If you self-host, your throughput is a function of your own hardware and serving stack, and the tier distinction stops existing.

The other thing worth flagging is that “roughly 180 tokens per second” is Moonshot’s own figure, published by the vendor and not independently verified by us. Treat it as a vendor claim until someone measures it — including us. It is on our list.

FAQ

Kimi K2.7 Code HighSpeed questions

What is Kimi K2.7 Code HighSpeed?

Kimi K2.7 Code HighSpeed is the same model as Kimi K2.7 Code served at roughly 180 tokens per second, at double the price.

What is the Kimi K2.7 Code HighSpeed context window?

Kimi K2.7 Code HighSpeed has a 262,144-token context window. That total covers input and output combined, so a long prompt reduces the room left for the response.

How much does Kimi K2.7 Code HighSpeed cost?

Kimi K2.7 Code HighSpeed costs $1.90 per million input tokens on a cache miss, $0.38 per million on a cache hit, and $8.00 per million output tokens. Prices exclude tax.

Is Kimi K2.7 Code HighSpeed open source?

Moonshot AI published downloadable open weights for Kimi K2.7 Code HighSpeed, so you can self-host it. That is not the same as an OSI-approved open-source licence — check the licence attached to the weight release before deploying commercially.

Is Kimi K2.7 Code HighSpeed a better model than Kimi K2.7 Code?

No. They are the same weights and produce the same quality of output. HighSpeed is a faster serving tier at exactly double the price. If output quality is what you are shopping for, the two are indistinguishable and you should use the cheaper one.

How fast is Kimi K2.7 Code HighSpeed?

Moonshot AI quotes roughly 180 tokens per second, reaching up to 260 tokens per second in short-context scenarios. Moonshot does not publish a throughput figure for standard Kimi K2.7 Code, so the size of the speed-up is not something we can state precisely.

When is HighSpeed worth double the price?

When a person is watching the output arrive. Interactive editor completions and live chat are latency-bound and benefit. Batch jobs, CI runs, nightly agents and anything queued are not, and there the markup buys nothing measurable.

Sources

Where these figures come from

Every specification and price on this page is transcribed from Moonshot AI’s own documentation and re-checked on the date shown. Nothing is estimated. Our editorial policy sets out how we source figures and what we do when something cannot be verified.