Kimi K3
Moonshot AI · released Jun 13, 2026 · moonshotai/Kimi-K3
- Type
- Open weightsCustom licence
- Params
- 2.8T
- Context
- 1M
104B active per word · about 786K words of context · download allowed, licence restricts use
Our take
Written Sep 1, 2026Kimi K3 is Moonshot AI's massive-scale mixture-of-experts model with 2.8 trillion total parameters and 104 billion active per word. It excels at reasoning and web-development tasks, and handles up to one million tokens in a single request, though its weights carry commercial restrictions.
Choose this for frontier reasoning where benchmark scores matter, web-application building where it leads on arena tasks, or long-document analysis at one-million-token scale. Pick it if you want provider choice: 30 offers create real price competition. Skip it if you need a permissive licence, agentic coding, or budget-tier pricing.
The case for it
- LiveBench Reasoning 90.67% and Mathematics 84.44% on monthly-refreshed tasks.
- Arena Code (WebDev) 1673.7, well ahead of its own general coding and text scores.
- One-million-token request limit, four times the 256K that was until recently considered frontier.
- Thirty current hosted offers, with the fastest provider at 59 tokens per second.
The case against it
- No budget tier among 30 offers; output pricing sits well above mid-size alternatives.
- Agentic coding lags its own coding score by 19.3 points: LiveBench Agentic Coding 62.17% versus LiveBench Coding 81.45%.
- Custom licence with commercial restrictions — not Apache, MIT, or equivalent.
How good is it?
An open-weights text model for coding, everyday questions, writing and multi-step tool work.
- getting answers to everyday questionsArena Text (overall) · 12th of 168
- drafts, rewrites and editingArena Creative Writing · 18th of 168
- writing and completing codeArena Coding · 6th of 168
- multi-step work it carries out for youArena Agent · 10th of 55
- calling tools to carry out requestsArena Agent · Tool use · 9th of 55
EverydayGeneral questions and everyday reasoning
Arena Text (overall)12th of 168 · 1488
CodingWriting and fixing code on its own
Arena Coding6th of 168 · 1541
AgenticPlanning, calling tools, staying on task
Arena Agent10th of 55 · 0.046
WritingDrafting and rewriting prose
Arena Creative Writing18th of 168 · 1458
Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.
Boards this model appears on that none of the ratings above are built on.
Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Apple M1 Pro (16-core GPU) · 32 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Radeon RX 7900 XT · 20 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 4 hours and 26 hours ago — each listing carries its own date.
Makora, through OpenRouter
Cheapest of the 11 listings we can compare like for like — at 1M of context, out of 24 in the table below. 2 cheaper rows there are outside that comparison: a different quantisation.
- per 1M tokens
- $2.04 in / $12.75 out
- Context served
- 1M
- Throughput
- ~33 tok/s
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| Morphfp8Through OpenRouter | $2.18 / $9.81checked 10 hours ago | 1M944K max reply | 52 tok/s | No | No | Confirmed |
| Relacefp4Through OpenRouter | $0.67 / $10.00checked 4 hours ago | 1M944K max reply | 31 tok/s | No | No | Confirmed |
| PhalaThrough OpenRouter | $2.55 / $12.75checked 4 hours ago | 1M944K max reply | 26 tok/s | No | No | Confirmed |
| MakoraThrough OpenRouter | $2.04 / $12.75checked 4 hours ago | 1M944K max reply | 33 tok/s | No | No | Confirmed |
| DigitalOcean GradientThrough OpenRouter | $2.55 / $12.95checked 4 hours ago | 1M944K max reply | 37 tok/s | No | No | Confirmed |
| InferenceNetfp4Through OpenRouter | $1.19 / $13.00checked 4 hours ago | 1M944K max reply | 64 tok/s | No | No | Confirmed |
| Together AIThrough OpenRouter | $2.70 / $13.50checked 4 hours ago | 1M944K max reply | 50 tok/s | No | No | Confirmed |
| Sail Researchfp4Through OpenRouter | $2.80 / $14.00checked 4 hours ago | 1M944K max reply | 42 tok/s | No | No | Confirmed |
| WaferusThrough OpenRouter | $2.80 / $14.00checked 4 hours ago | 1M944K max reply | 35 tok/s | No | No | Confirmed |
| WaferThrough OpenRouter | $1.39 / $14.00checked 4 hours ago | 1M944K max reply | 36 tok/s | No | No | Confirmed |
| DeepInfrabf16Direct and through OpenRouter | $2.85 / $14.25checked 4 hours ago | 1M16K max reply through OpenRouter | 9 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| AkashMLfp4Through OpenRouter | $3.00 / $15.00checked 4 hours ago | 1M944K max reply | 38 tok/s | No | No | Confirmed |
| Novita AIDirect | $3.00 / $15.00checked 4 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| Chutesmxfp4Through OpenRouter | $3.00 / $15.00checked 4 hours ago | 1M66K max reply | 15 tok/s | No | Yesunknown period | Unknown |
| Parasailfp4Through OpenRouter | $3.00 / $15.00checked 4 hours ago | 1M944K max reply | 72 tok/s | No | No | Confirmed |
| Basetenfp8Through OpenRouter | $3.00 / $15.00checked 4 hours ago | 1M262K max reply | 67 tok/s | No | No | Confirmed |
| OpenRouterOpenRouter's own listing | $3.00 / $15.00checked 26 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| Moonshot AImxfp4Through OpenRouter | $3.00 / $15.00checked 4 hours ago | 1M944K max reply | 24 tok/s | No | No | Confirmed |
| Fireworks AIThrough OpenRouter | $3.00 / $15.00checked 4 hours ago | 1M944K max reply | 21 tok/s | No | No | Confirmed |
| Decartmxfp4Through OpenRouter | $3.00 / $15.00checked 4 hours ago | 1M944K max reply | 40 tok/s | No | No | Confirmed |
| Modalmxfp4Through OpenRouter | $3.00 / $15.00checked 4 hours ago | 1M944K max reply | 76 tok/s | No | No | Confirmed |
| Alibaba CloudThrough OpenRouter | $3.45 / $17.25checked 4 hours ago | 1M944K max reply | 42 tok/s | No | Yesunknown period | Unknown |
| Fireworks AIfast tierThrough OpenRouter | $4.50 / $22.50checked 4 hours ago | 1M944K max reply | 64 tok/s | No | No | Confirmed |
| Fireworks AIusThrough OpenRouter | $4.50 / $22.50checked 4 hours ago | 1M944K max reply | 65 tok/s | No | No | Confirmed |
Across the 24 listings we hold: 22 say they do not train on prompts (1 of them only through OpenRouter), 0 say they do and 2 do not say. 20 appear in the zero-retention registry we check (1 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| Morphfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Relacefp4Through OpenRouter | ✓ | ✗ | ✗ |
| PhalaThrough OpenRouter | ✓ | ✓ | ✓ |
| MakoraThrough OpenRouter | ✓ | ✓ | ✓ |
| DigitalOcean GradientThrough OpenRouter | ✓ | ✓ | ✗ |
| InferenceNetfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Together AIThrough OpenRouter | ✓ | ✓ | ✓ |
| Sail Researchfp4Through OpenRouter | ✓ | ✓ | ✓ |
| WaferusThrough OpenRouter | ✓ | ✓ | ✓ |
| WaferThrough OpenRouter | ✓ | ✓ | ✓ |
| DeepInfrabf16Direct and through OpenRouter | ✓ | ✓ | ✓ |
| AkashMLfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Novita AIDirect | |||
| Chutesmxfp4Through OpenRouter | ✗ | ✓ | ✓ |
| Parasailfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Basetenfp8Through OpenRouter | ✓ | ✗ | ✗ |
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| Moonshot AImxfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Fireworks AIThrough OpenRouter | ✓ | ✓ | ✓ |
| Decartmxfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Modalmxfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Alibaba CloudThrough OpenRouter | ✓ | ✓ | ✓ |
| Fireworks AIfastThrough OpenRouter | ✗ | ✓ | ✓ |
| Fireworks AIusThrough OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 21 of 24 listings say yes, 2 say no, 1 publishes no parameter list. JSON output: 21 of 24 listings say yes, 2 say no, 1 publishes no parameter list. Strict schema: 20 of 24 listings say yes, 3 say no, 1 publishes no parameter list.
Models people weigh against Kimi K3
When we formed this view
Recent changes
What moved
Kimi K3 moved on 6 hosts: Sail Research: input +781% ($0.318 → $2.800 per 1M tokens), output +76% ($7.94 → $14.00 per 1M tokens), cache read −25% ($0.40 → $0.30 per 1M tokens); InferenceNet: input +526% ($0.19 → $1.19 per 1M tokens), output +18% ($11.00 → $13.00 per 1M tokens), cache read +344% ($0.18 → $0.80 per 1M tokens); Morph: input +86% ($1.17 → $2.18 per 1M tokens), output −28% ($13.55 → $9.81 per 1M tokens); Relace: input −68% ($2.10 → $0.67 per 1M tokens), output −5% ($10.50 → $10.00 per 1M tokens), cache read +219% ($0.21 → $0.67 per 1M tokens); Wafer: input −4% ($1.442 → $1.390 per 1M tokens), output +56% ($9.00 → $14.00 per 1M tokens); Together: input −10% ($3.00 → $2.70 per 1M tokens), output −10% ($15.00 → $13.50 per 1M tokens), cache read −10% ($0.30 → $0.27 per 1M tokens); Wafer (US region): cache read +7% ($0.28 → $0.30 per 1M tokens)What moved
Kimi K3 moved on 4 hosts: Sail Research: input −61% ($0.809 → $0.318 per 1M tokens), output −38% ($12.75 → $7.94 per 1M tokens); InferenceNet: input −53% ($0.40 → $0.19 per 1M tokens), output +22% ($9.00 → $11.00 per 1M tokens), cache read −55% ($0.40 → $0.18 per 1M tokens); Fireworks (US region): input +36% ($3.30 → $4.50 per 1M tokens), output +36% ($16.50 → $22.50 per 1M tokens), cache read +36% ($0.33 → $0.45 per 1M tokens); Morph: input −5% ($1.23 → $1.17 per 1M tokens), output +27% ($10.70 → $13.55 per 1M tokens)What moved
Kimi K3 moved on 4 hosts: Sail Research: input −10% ($0.90 → $0.81 per 1M tokens), output +62% ($7.86 → $12.75 per 1M tokens), cache read +33% ($0.30 → $0.40 per 1M tokens); InferenceNet: input −60% ($1.00 → $0.40 per 1M tokens), cache read +33% ($0.30 → $0.40 per 1M tokens); Wafer: input +44% ($1.00 → $1.44 per 1M tokens); Morph: input +40% ($0.88 → $1.23 per 1M tokens), output −4% ($11.20 → $10.70 per 1M tokens)What moved
Kimi K3 moved on 2 hosts: Morph: input −65% ($2.50 → $0.88 per 1M tokens), output −20% ($14.00 → $11.20 per 1M tokens); Sail Research: input −13% ($1.03 → $0.90 per 1M tokens), output −13% ($9.04 → $7.86 per 1M tokens)What moved
input +19% ($1.00 → $1.19 per 1M tokens)What moved
Kimi K3 moved on 2 hosts: InferenceNet: input −23% ($1.30 → $1.00 per 1M tokens), output −21% ($11.40 → $9.00 per 1M tokens); Sail Research: input −14% ($1.20 → $1.03 per 1M tokens), output −14% ($10.53 → $9.04 per 1M tokens)What moved
Kimi K3 moved on 4 hosts: Makora: input −20% ($2.55 → $2.04 per 1M tokens), cache read −20% ($0.256 → $0.204 per 1M tokens); Sail Research: input −11% ($1.35 → $1.20 per 1M tokens), output −8% ($11.50 → $10.53 per 1M tokens); InferenceNet: input −7% ($1.40 → $1.30 per 1M tokens), output +6% ($10.75 → $11.40 per 1M tokens); Wafer: output +3% ($14.50 → $15.00 per 1M tokens), cache read +51% ($0.199 → $0.300 per 1M tokens)Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- 1 of 24 listings publishes no parameter list, so what its API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 2 of 24 listings do not say whether they train on prompts, and 1 answers only through OpenRouter, not for its own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsCustom licence, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
Custom licence
This model ships custom license terms that don't map to a known template. We haven't parsed them, so commercial use, redistribution and derivatives are unverified — review the original terms before shipping.
Identifiers
- Hugging Face
- moonshotai/Kimi-K3
- Architecture
- Mixture of experts
- Takes in, gives back
- Text and images in, text out
- Catalogue slug
- moonshotai-kimi-k3