Grok 4.20
xAI · released Mar 31, 2026
- Type
- Closed
- Input
- $1.25
- Output
- $2.50
- Cached
- None held
List price · per 1M tokens · xAI at 1M context · machine-readable source ↗
Our take
Written Sep 3, 2026Grok 4.20 is xAI's flagship text model with a two-million-token request limit and measured coding performance that leads its own scorecard. It is a premium hosted-only option for long-document and coding work where you pay for throughput if speed matters.
Choose this for coding tasks where its measured coding score is the relevant signal, or for long-document analysis across millions of tokens. Pick the premium tier if 91 tokens per second is worth twice the standard rate to your workflow. Skip it if you need open weights, if your budget rules out premium throughput, or if creative writing and instruction following are the main task.
The case for it
- Two-million-token request limit — ten times the length common among frontier models we track.
- Strong measured coding performance, 33.6 points above its own overall text score.
- Six Arena variants measured, all scoring in the mid-1400s to low-1500s range.
- Throughput option up to 91 tokens per second on the premium tier.
The case against it
- Premium throughput costs double the standard rate for roughly 14% more speed.
- Proprietary weights with no open licence or disclosed parameter count.
- Creative writing and instruction following lag its own coding score by 46 or more points.
How good is it?
A closed text model from xAI for everyday questions, drafting and coding.
- getting answers to everyday questionsArena Text (overall) · 25th of 168
- drafts, rewrites and editingArena Creative Writing · 15th of 168
- writing and completing codeArena Coding · 40th of 168
EverydayGeneral questions and everyday reasoning
Arena Text (overall)25th of 168 · 1475
CodingWriting and fixing code on its own
Arena Coding40th of 168 · 1508
Arena Coding is the only board that has scored it for this.
AgenticPlanning, calling tools, staying on task
Not yet scored on Arena Agent.
WritingDrafting and rewriting prose
Arena Creative Writing15th of 168 · 1464
Arena Creative Writing is the only board that has scored it for this.
These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done.
Every published score for this model6 scoresEvery figure we hold, from 6 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Where to rent it
Prices checked 4 hours ago — each listing carries its own date.
xAI, through OpenRouter
Why this differs from the header. The strip above quotes xAI's own list price; this is xAI through OpenRouter, which lists it at a lower price.
- per 1M tokens
- $1.25 in / $2.50 out
- Context served
- 2M
- Throughput
- ~57 tok/s
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| OpenRouterOpenRouter's own listing | $1.25 / $2.50checked 4 hours ago | 2M | not measured | Unknown | Unknown | Unknown |
| xAIDirect | $1.25 / $2.50checked 4 hours ago | 1M1M max reply | not measured | Unknown | Unknown | Unknown |
| xAIThrough OpenRouter | $1.25 / $2.50checked 4 hours ago | 2M1.8M max reply | 57 tok/s | No | Yes30 days | Confirmed |
| xAIpriority tierThrough OpenRouter | $2.50 / $5.00checked 4 hours ago | 2M1.8M max reply | 84 tok/s | No | Yes30 days | Confirmed |
Across the 4 listings we hold: 2 say they do not train on prompts, 0 say they do and 2 do not say. 2 appear in the zero-retention registry we check; the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| xAIDirect | |||
| xAIThrough OpenRouter | ✓ | ✓ | ✓ |
| xAIpriorityThrough OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 3 of 4 listings say yes, 1 publishes no parameter list. JSON output: 3 of 4 listings say yes, 1 publishes no parameter list. Strict schema: 3 of 4 listings say yes, 1 publishes no parameter list.
Models people weigh against Grok 4.20
When we formed this view
Recent changes
What moved
first indexed by our pipelineEach date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- 1 of 4 listings publishes no parameter list, so what its API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 2 of 4 listings do not say whether they train on prompts.
- We hold no batch or off-peak rate for any of its listings.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.
Identifiers
- Takes in, gives back
- Text, images and documents in, text out
- Catalogue slug
- x-ai-grok-4-20