Official comparison

Kimi K3 vs K2.7 Code: Official Specifications

Side-by-side official specs and pricing for Kimi K3 and Kimi K2.7 Code, focused on API model IDs, context windows, thinking controls, tool behavior, structured output, caching, and CNY pricing.

As of July 24, 2026

Quick conclusion

Use the task constraints before choosing a model

Kimi K3 has the larger official context window and supports tool_choice: required. Kimi K2.7 Code has lower official CNY per-token pricing in the accessed sources, while Kimi K2.7 Code USD pricing was not stated in the official sources accessed as of July 24, 2026.

Official specifications

Quick comparison table

These fields come from audited official Kimi and Moonshot AI sources. API parameters and pricing may change over time.

Official specification comparison for Kimi K3 and Kimi K2.7 Code
FieldKimi K3Kimi K2.7 CodeSources
Official positioningMoonshot AI positions Kimi K3 as its flagship model for long-horizon coding, knowledge work, and deep reasoning.Kimi K2.7 Code is positioned as a specialized coding model.CS1, CS3
API model IDkimi-k3kimi-k2.7-code; high-speed variant: kimi-k2.7-code-highspeedCS2, CS5, CS3
Model list observationAs of July 24, 2026, kimi-k3 appeared in the official Kimi API model list.As of July 24, 2026, kimi-k2.7-code appeared in the official Kimi API model list.CS5
Context window1,048,576 tokens (1M).262,144 tokens (256K).CS2, CS4, CS3, CS7
Thinking/reasoning controlThinking is always on. Kimi K3 supports reasoning_effort: low, high, or max; default is max.Thinking is always on. Kimi K2.7 Code uses thinking: {"type": "enabled"}; no intensity control is documented.CS2, CS4, CS3
Output token settingDefault max output tokens: 131,072. Configurable up to 1,048,576.Default max output tokens: 32,768. Maximum configurable value was not confirmed in the accessed official sources as of July 24, 2026.CS2, CS3
tool_choiceSupports auto, none, and required.Supports auto and none only; required is not supported and returns an error.CS2, CS4, CS3
Dynamic tool loadingSupported and documented in the K3 quickstart with a code example.Not documented in the K2.7 Code quickstart as of July 24, 2026; absence from that guide does not confirm unsupported status.CS2, CS3
Structured outputSupports response_format with json_schema and strict: true.The K2.7 Code pricing page lists JSON Mode; detailed interface documentation was not found in the K2.7 Code quickstart as of July 24, 2026.CS2, CS7
Partial ModeSupported and documented in the K3 quickstart with a code example.Listed in the K2.7 Code pricing page feature summary; detailed usage documentation was not found in the quickstart as of July 24, 2026.CS2, CS7
Context cachingAutomatic. Moonshot AI reports a cache hit rate above 90% in coding workloads.Automatic, stated in the pricing page feature summary. Cache hit rate was not stated in the official sources accessed as of July 24, 2026.CS1, CS2, CS7

Official pricing

CNY pricing per 1M tokens

Pricing is highly time-sensitive. The values below are from official sources accessed as of July 24, 2026; prices may change.

CNY pricing comparison for Kimi K3 and Kimi K2.7 Code
Pricing fieldKimi K3Kimi K2.7 Code
CNY per 1M tokensPublished in CNY on the official K3 pricing page.Published in CNY on the official K2.7 Code pricing page.
Cache hit input¥2.00¥1.30 standard; ¥2.60 Highspeed
Cache miss input¥20.00¥6.50 standard; ¥13.00 Highspeed
Output¥100.00¥27.00 standard; ¥54.00 Highspeed
USD pricingMoonshot AI's K3 announcement blog states approximately $0.30 cache hit input, $3.00 cache miss input, and $15.00 output per 1M tokens.Not stated in the official sources accessed as of July 24, 2026.
Prices may changeAs of July 24, 2026. Prices may change. Check the current official pricing pages for the latest rates.As of July 24, 2026. Prices may change. Check the current official pricing pages for the latest rates.
Official pricing sourceCS6 for CNY pricing; CS8 for the blog USD statement.CS7 for CNY pricing.

Conditional use cases

When each official difference may matter

These are conditional suggestions based on official specifications, vendor-reported speed, tool_choice support, and official CNY unit prices.

If a task requires processing very long inputs, such as large codebases or lengthy documents, Kimi K3's 1M token context window may be advantageous compared to K2.7 Code's 256K window, based on official specifications.

According to Moonshot AI's published specifications, Kimi K2.7 Code Highspeed outputs approximately 180 tokens per second, reaching up to 260 tokens per second in short-context scenarios. If coding task throughput is the primary consideration, this vendor-stated speed may be relevant. Independent speed verification has not been performed.

Kimi K3 supports tool_choice: required, which guarantees the model will invoke at least one tool in the response. Kimi K2.7 Code supports tool_choice values of auto and none only. Based on these documented API differences, K3 may be more suitable for agentic workflows that require guaranteed tool execution. This has not been verified through comparative testing.

Based on official pricing pages accessed July 24, 2026, Kimi K2.7 Code has lower per-token pricing than Kimi K3 across all billing categories: cache hit, cache miss, and output. Actual cost depends on cache hit rates, prompt length, and task patterns. Kimi K2.7 Code USD pricing was not stated in the official sources accessed as of July 24, 2026.

Original experiment

Independent tests have not been executed

The planned evaluation will cover four task types: TypeScript bug fix, React component from specification, Screenshot restoration, Multi-file refactoring.

Scope and limitations

What this page does not claim

This page does not publish search volume, K2.7 Code parameter counts or architecture, side-by-side benchmark conclusions, user sentiment, independent result data, or a claim that K3 has replaced K2.7 Code.

Official sources

Kimi and Moonshot AI source documents

The comparison above is limited to the official sources listed here and accessed as of July 24, 2026.

FAQ

Questions developers ask about Kimi K3 and K2.7 Code

What is the main difference between Kimi K3 and Kimi K2.7 Code?

Based on official sources accessed as of July 24, 2026, Moonshot AI positions Kimi K3 as a flagship model for long-horizon coding, knowledge work, and deep reasoning, while Kimi K2.7 Code is positioned as a specialized coding model.

Which model has the larger context window?

Kimi K3 has the larger stated context window at 1,048,576 tokens, compared with 262,144 tokens for Kimi K2.7 Code, based on official specifications accessed as of July 24, 2026.

Which has lower official per-token pricing?

Based on official CNY pricing pages accessed July 24, 2026, Kimi K2.7 Code has lower per-token pricing across cache hit input, cache miss input, and output. Actual cost may vary by cache hit rate, prompt length, and task pattern.

Does K3 replace K2.7 Code?

No official source accessed as of July 24, 2026 says Kimi K3 replaces Kimi K2.7 Code. Both kimi-k3 and kimi-k2.7-code appeared in the official Kimi API model list on the access date.

Have you independently tested the models?

No. Original experiments have not been executed, so this page does not publish independent result data or make a comparative outcome claim.