Models / xAI/ Grok Build 0.1

Grok Build 0.1

xAI · released May 20, 2026

Input: text, images and documents. Output: text.InputOutput
Type
Closed
Input
$1.00
Output
$2.00
Cached
None held

List price · per 1M tokens · xAI at 256K context · machine-readable source ↗

Our take

Written Sep 30, 2026

Grok Build 0.1 is a hosted-only model with one clear speciality: picking the right tool and not inventing one, where it places 8th of 55 on Arena Agent · Tool use as of 25 Sep 2026. Everywhere else its measured results sit near the bottom of the field, so it is a narrow pick rather than a general one.

Who should pick it

Reach for it when the job is agent work that hinges on choosing the right tool, or when long documents would otherwise have to be split before sending. We list no download for it, so using it means choosing a host. Skip it if you need code written or maths solved, or if you need an agent to recover after a command fails or hold a course correction.

The case for it

  • Tool selection is the one thing it does well: 8th of 55 on Arena Agent · Tool use as of 25 Sep 2026, an inverse-propensity score for calling the right tool and not inventing one, far ahead of its own 49th of 55 on the parent Arena Agent board.
  • Long inputs need not be split up first, though whether it recalls everything inside them is unverified in our data.
  • Files and images go in with the question, so a document or a screenshot does not have to be retyped as prose.

The case against it

  • Coding is among the weakest measured areas: 58th of 58 on LiveBench Coding as of 25 Jun 2026, a percentage average over code generation and completion tasks.
  • Maths and reasoning sit near the bottom of the same field: 55th of 58 on LiveBench Mathematics and 50th of 58 on LiveBench Reasoning as of 25 Jun 2026.
  • Agent sessions that go wrong tend not to come back: 53rd of 55 on Arena Agent · Recovery as of 25 Sep 2026, and 41st of 55 on task outcome.
00

How good is it?

A text model for calling tools to carry out requests, though multi-step work and changing course are weaker spots.

Good at
  • calling tools to carry out requestsArena Agent · Tool use · 8th of 55
Less good at
  • multi-step work it carries out for youArena Agent · 49th of 55
  • changing course when you give new instructionsArena Agent · Steerability · 45th of 55
  • getting back on track after a step failsArena Agent · Recovery · 53rd of 55

EverydayGeneral questions and everyday reasoning

Scored, not ratedLiveBench Data Analysis · 41st of 58 · 70.79

Not yet scored on Arena Text (overall). It is on LiveBench Data Analysis, in 41st of 58 with 70.79.

LiveBench Reasoning 50th of 58LiveBench Mathematics 55th of 58

CodingWriting and fixing code on its own

Scored, not ratedLiveBench Coding · 58th of 58 · 65.39

Not yet scored on Arena Coding. It is on LiveBench Coding, in 58th of 58 with 65.39.

AgenticPlanning, calling tools, staying on task

1 of 5

Arena Agent49th of 55 · −0.129

LiveBench Agentic Coding 44th of 58

WritingDrafting and rewriting prose

Scored, not ratedLiveBench Language · 51st of 58 · 72.46

Not yet scored on Arena Creative Writing. It is on LiveBench Language, in 51st of 58 with 72.46.

How it behaves in an agent loop
Tool usereaches for the right one, and does not invent one8th of 55
Steerabilitydoes what it was asked, and changes course when told45th of 55
Recoverygets back on track after a command fails53rd of 55
Task outcomefinishes what the session set out to do41st of 55

Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.

Other boards it appears on
LiveBench Instruction Following 36th of 58LiveBench 51st of 58Arena Agent · Tool use 8th of 55Arena Agent · Task outcome 41st of 55Arena Agent · Steerability 45th of 55Arena Agent · Recovery 53rd of 55

Boards this model appears on that none of the ratings above are built on.

Every published score for this model13 scoresEvery figure we hold, from 13 boards, with who ran it and a link to the source — including the boards no rating above is built on.
LiveBenchreasoning
67.78source ↗
45.81source ↗
65.39source ↗
70.79source ↗
72.46source ↗
78.43source ↗
76.37source ↗
−0.129source ↗
−0.392source ↗
−0.067source ↗
−0.074source ↗
0.005source ↗
01

Where to rent it

Prices checked 4 hours ago — each listing carries its own date.

Cheapest published offer

xAI, direct and through OpenRouter

The lab is also the cheapest we hold. The strip above and this offer are the same one, so nothing on this page undercuts xAI on 256K of context.

per 1M tokens
$1.00 in / $2.00 out
Context served
256K
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
OpenRouterOpenRouter's own listing$1.00 / $2.00checked 4 hours ago256Knot measuredUnknownUnknownUnknown
xAIDirect and through OpenRouter$1.00 / $2.00checked 4 hours ago256K256K max reply direct230K max reply through OpenRouter101 tok/sthrough OpenRouterDirectUnknownThrough OpenRouterNoDirectUnknownThrough OpenRouterYes30 daysDirectUnknownThrough OpenRouterConfirmed
xAIpriority tierThrough OpenRouter$2.00 / $4.00checked 4 hours ago256K230K max reply73 tok/sNoYes30 daysConfirmed

Across the 3 listings we hold: 2 say they do not train on prompts (1 of them only through OpenRouter), 0 say they do and 1 does not say. 2 appear in the zero-retention registry we check (1 of them only through OpenRouter); the rest are unknown to us.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

API features per host
ProviderTool callingJSON outputStrict schema
OpenRouterOpenRouter's own listing✓✓✓
xAIDirect and through OpenRouter✓✓✓
xAIpriorityThrough OpenRouter✓✓✓

Tool calling: 3 of 3 listings say yes. JSON output: 3 of 3 listings say yes. Strict schema: 3 of 3 listings say yes.

02

When we formed this view

Recent changes

Sep 5, 2026BenchmarkScored −0.129 on Arena Agent
What movedleaderboard
Sep 5, 2026BenchmarkScored −0.392 on Arena Agent · Recovery
What movedleaderboard
Sep 5, 2026BenchmarkScored −0.067 on Arena Agent · Steerability
What movedleaderboard
Sep 5, 2026BenchmarkScored −0.074 on Arena Agent · Task outcome
What movedleaderboard
Sep 5, 2026BenchmarkScored 0.005 on Arena Agent · Tool use
What movedleaderboard
Jul 26, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline
Jun 25, 2026BenchmarkScored 67.78 on LiveBench
What movedleaderboard
Jun 25, 2026BenchmarkScored 45.81 on LiveBench Agentic Coding
What movedleaderboard
Jun 25, 2026BenchmarkScored 65.39 on LiveBench Coding
What movedleaderboard
Jun 25, 2026BenchmarkScored 70.79 on LiveBench Data Analysis
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 3 listings does not say whether it trains on prompts, and 1 answers only through OpenRouter, not for its own listing.
  • We hold no batch or off-peak rate for any of its listings.
03

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Text, images and documents in, text out
Catalogue slug
x-ai-grok-build-0-1

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us