Kimi K2 Thinking
Moonshot AI · released Nov 4, 2025 · moonshotai/Kimi-K2-Thinking
- Type
- Open weightsCustom licence
- Params
- 1.1T
- Context
- 262K
32B active per word · about 197K words of context · download allowed, licence restricts use
Our take
Written Sep 16, 2026Kimi K2 Thinking is a large text-only model from Moonshot AI built for coding and hard reasoning tasks, with 32 billion active parameters and a 262,144-token request limit. Its measured strengths sit in software engineering and code generation rather than creative writing.
Choose this for coding workflows where its Arena score peaks, or for long-context tasks needing a quarter-million tokens at a manageable active-parameter cost. Use it for end-to-end software engineering with a verified issue-resolution rate above sixty per cent. Skip it if you need a permissive licence, creative writing quality, or guaranteed fast throughput across every provider.
The case for it
- Coding is its standout skill, leading its own Arena Text score by over fifty points and sitting within thirty points of maths and hard prompts.
- Verified software engineering capability at 63.4% on SWE-bench Verified.
- 262,144-token request limit with only 32 billion active parameters — an unusually large window for the active compute cost.
- Identical pricing across all four tracked offers removes provider-hopping for cost.
The case against it
- Creative writing lags its coding score by nearly eighty points, the widest skill gap in its profile.
- Custom restricted licence — not Apache 2.0 or MIT — with redistribution and commercial terms limited.
- Throughput varies more than threefold between providers at the same price, and two plans disclose no speed data.
How good is it?
EverydayGeneral questions and everyday reasoning
Arena Text (overall)49th of 168 · 1450
CodingWriting and fixing code on its own
Arena Coding48th of 168 · 1502
AgenticPlanning, calling tools, staying on task
Not yet scored on Arena Agent. It is on SWE-bench Verified, in 21st of 42 with 63.4.
WritingDrafting and rewriting prose
Arena Creative Writing48th of 168 · 1423
Arena Creative Writing is the only board that has scored it for this.
These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done.
Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Apple M1 Pro (16-core GPU) · 32 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Radeon RX 7900 XT · 20 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked 4 hours ago — each listing carries its own date.
- per 1M tokens
- $0.60 in / $2.50 out
- Context served
- 262K
- Throughput
- ~106 tok/s
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| OpenRouterOpenRouter's own listing | $0.60 / $2.50checked 4 hours ago | 262K | not measured | Unknown | Unknown | Unknown |
| Novita AIbf16Direct and through OpenRouter | $0.60 / $2.50checked 4 hours ago | 262K98K max reply through OpenRouter | 47 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| Google Vertex AIThrough OpenRouter | $0.60 / $2.50checked 4 hours ago | 262K236K max reply | 106 tok/s | No | No | Confirmed |
Across the 3 listings we hold: 2 say they do not train on prompts (1 of them only through OpenRouter), 0 say they do and 1 does not say. 2 appear in the zero-retention registry we check (1 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| Novita AIbf16Direct and through OpenRouter | ✓ | ✓ | ✓ |
| Google Vertex AIThrough OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 3 of 3 listings say yes. JSON output: 3 of 3 listings say yes. Strict schema: 3 of 3 listings say yes.
When we formed this view
Recent changes
What moved
first indexed by our pipelineEach date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 1 of 3 listings does not say whether it trains on prompts, and 1 answers only through OpenRouter, not for its own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsCustom licence, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
Custom licence
This model ships custom license terms that don't map to a known template. We haven't parsed them, so commercial use, redistribution and derivatives are unverified — review the original terms before shipping.
Identifiers
- Hugging Face
- moonshotai/Kimi-K2-Thinking
- Architecture
- Mixture of experts
- Takes in, gives back
- Text in, text out
- Catalogue slug
- moonshotai-kimi-k2-thinking