Kimi K2.7 Code is a coding-oriented chat model from Kimi, positioned as the code-focused variant in the K2.7 line and available through inference API providers.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.680 | $3.40 | $0.136 | |
| $0.710 | $3.50 | $0.150 | |
| $0.872 | $3.83 | $0.436 | |
| $0.950 | $4.00 | - | |
| $0.950 | $4.00 | - | |
| $0.950 | $4.00 | - | |
| $0.950 | $4.00 | $0.190 | |
| $1.25 | $4.50 | $0.310 |
Prices updated daily. Last check: Sep 8, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Kimi K2.7 Code is aimed at development workloads: generating functions and modules from natural-language specifications, refactoring existing files, writing unit tests, explaining unfamiliar code, and diagnosing errors from stack traces. Its measured latency profile — roughly 1.3 seconds to first token with mid-range output speed — fits interactive use in an editor or terminal assistant, and it is workable inside agentic loops where a model reads code, proposes an edit, and reacts to test output, though each turn will feel deliberate rather than instant. For very high-volume, latency-critical tasks such as inline autocomplete, a smaller and faster model is usually the better fit, and for broad non-coding reasoning or writing, a general-purpose chat model is the more natural choice. Because we do not have a confirmed context window for this variant, teams planning whole-repository ingestion should confirm limits with their chosen provider before committing.
Pricing depends on which inference provider you use and how they bill — separate input and output token rates, cached input discounts, and batch or dedicated-capacity options all differ between hosts. Because those rates change frequently, we do not quote figures in this write-up; check the pricing table on this page for the current per-provider rates for Kimi K2.7 Code.
It is a code-focused variant, so it suits code generation, refactoring, test authoring, code explanation, and debugging assistance, including use as the model behind a coding agent or editor integration. Its moderate first-token latency and mid-range output speed make it more appropriate for deliberate assistant turns than for inline autocomplete.
Artificial Analysis measurements for this model show roughly 39.5 output tokens per second and about 1,280 ms to first token. Those are serving characteristics of a specific measured endpoint — actual speed varies by provider, prompt length, and load, so compare hosts if throughput matters to your workload.
The "Code" suffix marks it as the coding-tuned variant within the K2.7 line, as opposed to Kimi's general-purpose chat entries. If your workload is mostly software development, the code variant is the intended target; for mixed conversational, analytical, and writing tasks, a general Kimi model or another general-purpose chat model is usually a closer match.
We do not currently track a confirmed context window for this variant, and a missing value in our database means unverified rather than small. Check the documentation of the provider you plan to use, especially if you intend to feed in large codebases or long diffs.