Models / Z.ai/ GLM 5.3

GLM 5.3

Z.ai · released Aug 18, 2026

Input: text. Output: text.InputOutput
Type
Closed
Input
$1.40
Output
$4.40
Cached
$0.26

List price · per 1M tokens · Z.AI at 1M context · source ↗ · a reseller below undercuts it; the table carries the spread

Our take

Written Sep 30, 2026

GLM 5.3 is a hosted-only text model: we list no download for it, so using it means choosing a host. It is strongest where people rate the answers themselves, and mid-field on the harder task sets.

Who should pick it

Reach for it on chat, writing and general assistant work where human-rated answer quality matters, or in agent sessions that need reliable tool selection, where it places 9th of 55 on Arena Agent · Tool use via Max as of 25 Sep 2026. Long documents need not be split up first, though recall across all of it is unverified in our data. Skip it if you need to run the model on your own hardware, or if data analysis is the job.

The case for it

  • 16th of 168 on Arena Text (overall) via Max as of 25 Sep 2026, a board that records which answer people preferred rather than whether it was correct.
  • Among the better measured models at picking the right tool in an agent session: 9th of 55 on Arena Agent · Tool use via Max as of 25 Sep 2026, a board that scores calling the right tool and not inventing one.
  • 87.9% on LiveBench Mathematics and 85.8% on LiveBench Reasoning, both averages over competition-style and monthly-refreshed task sets rather than work in an existing project.
  • The request capacity takes a long report or a stack of documents alongside the question, so long inputs need not be split up first.

The case against it

  • We list no download for it, so every route we hold is a hosted offer, and the licence is not disclosed.
  • 44th of 58 on LiveBench Data Analysis as of 25 Jun 2026, a board of table and event-ordering tasks, so spreadsheet-style work is where it trails its own other results.
  • 60.91% on LiveBench Agentic Coding against 78.95% on LiveBench Coding, and the agentic figure was run inside an agent harness, so it reflects the model in that scaffold rather than on its own.
00

How good is it?

A closed text model for everyday questions, drafting, coding and calling tools.

Good at
  • getting answers to everyday questionsArena Text (overall) · 16th of 168
  • drafts, rewrites and editingArena Creative Writing · 19th of 168
  • writing and completing codeArena Coding · 23rd of 168
  • calling tools to carry out requestsArena Agent · Tool use · 9th of 55

EverydayGeneral questions and everyday reasoning

4 of 5

Arena Text (overall)16th of 168 · 1480

Arena Hard Prompts 15th of 168Arena Maths 13th of 163LiveBench Reasoning 28th of 58LiveBench Mathematics 37th of 58LiveBench Data Analysis 44th of 58

CodingWriting and fixing code on its own

4 of 5

Arena Coding23rd of 168 · 1522

Arena Code (WebDev) 15th of 95LiveBench Coding 22nd of 58

AgenticPlanning, calling tools, staying on task

3 of 5

Arena Agent17th of 55 · 0.029

LiveBench Agentic Coding 13th of 58

WritingDrafting and rewriting prose

3.5 of 5

Arena Creative Writing19th of 168 · 1456

LiveBench Language 26th of 58
How it behaves in an agent loop
Tool usereaches for the right one, and does not invent one9th of 55
Steerabilitydoes what it was asked, and changes course when told17th of 55
Recoverygets back on track after a command fails32nd of 55
Task outcomefinishes what the session set out to do13th of 55

Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.

Other boards it appears on
Arena Instruction Following 12th of 168LiveBench 25th of 58LiveBench Instruction Following 29th of 58Arena Agent · Tool use 9th of 55Arena Agent · Task outcome 13th of 55Arena Agent · Steerability 17th of 55Arena Agent · Recovery 32nd of 55

Boards this model appears on that none of the ratings above are built on.

Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
LiveBenchreasoning
76.14source ↗
60.91source ↗
78.95source ↗
70.24source ↗
79.86source ↗
87.9source ↗
85.8source ↗
0.029source ↗
−0.008source ↗
0.03source ↗
0.061source ↗
0.004source ↗
1522source ↗
1456source ↗
1506source ↗
1480source ↗
1497source ↗
1480source ↗
1619source ↗
01

Where to rent it

Prices checked between 3 hours and 4 days ago — each listing carries its own date.

Some hosts sell this model at two prices: on their own price list (“direct”) and on their OpenRouter listing (“through OpenRouter”). Where the two differ, the row shows both, each with the date we last read it.

Cheapest published offer

InferenceNet, through OpenRouter

Why this differs from the header. The strip above quotes Z.AI's own list price; this is the cheapest live offer, whoever is serving it — a reseller undercutting a lab is ordinary commerce, not an error.

per 1M tokens
$0.68 in / $2.28 out
Context served
1M
Throughput
~91 tok/s
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
RekaThrough OpenRouter$0.37 / $1.14checked 9 hours ago262K236K max reply60 tok/sNoNoConfirmed
SiliconFlowfp8Through OpenRouter$0.70 / $2.20checked 3 hours ago1M262K max reply54 tok/sNoNoConfirmed
InferenceNetThrough OpenRouter$0.68 / $2.28checked 2 days ago1M131K max reply91 tok/sNoNoConfirmed
Novita AIfp8Direct and through OpenRouter$1.40 / $4.40directchecked 3 hours ago$0.78 / $2.46through OpenRouterchecked 3 hours ago1M131K max reply through OpenRouter45 tok/sthrough OpenRouterDirectUnknownThrough OpenRouterNoDirectUnknownThrough OpenRouterNoDirectUnknownThrough OpenRouterConfirmed
DeepInfrafp4Direct and through OpenRouter$0.90 / $4.00directchecked 3 hours ago$0.56 / $2.50through OpenRouterchecked 3 hours ago1M131K max reply through OpenRouter36 tok/sthrough OpenRouterDirectUnknownThrough OpenRouterNoDirectUnknownThrough OpenRouterNoDirectUnknownThrough OpenRouterConfirmed
PhalaThrough OpenRouter$0.84 / $2.64checked 3 hours ago1M944K max reply44 tok/sNoNoConfirmed
DigitalOcean GradientThrough OpenRouter$0.91 / $2.86checked 3 hours ago1M128K max reply47 tok/sNoNoConfirmed
Morphfp8Through OpenRouter$1.19 / $2.99checked 2 days ago1M944K max reply109 tok/sNoNoConfirmed
GMICloudfp8Through OpenRouter$0.98 / $3.08checked 3 hours ago1M944K max reply53 tok/sNoYesunknown periodUnknown
Inceptronfp4Through OpenRouter$0.60 / $3.39checked 3 hours ago1M944K max reply60 tok/sNoNoConfirmed
AkashMLfp8Through OpenRouter$1.05 / $3.56checked 3 hours ago1M131K max reply80 tok/sNoNoConfirmed
Decartfp4Through OpenRouter$1.19 / $3.74checked 3 hours ago1M944K max reply129 tok/sNoNoConfirmed
Alibaba CloudThrough OpenRouter$1.19 / $3.74checked 3 hours ago1M131K max reply74 tok/sNoYesunknown periodUnknown
Makorafp4Through OpenRouter$0.85 / $3.93checked 9 hours ago980K128K max reply109 tok/sNoNoConfirmed
FriendliThrough OpenRouter$1.26 / $3.96checked 3 hours ago1M944K max reply92 tok/sNoYesunknown periodUnknown
Sail Researchusfp8Through OpenRouter$0.77 / $4.00checked 3 days ago1M944K max reply83 tok/sNoNoUnknown
Sail Researchfp8Through OpenRouter$0.77 / $4.00checked 3 days ago1M944K max reply74 tok/sNoNoConfirmed
RelaceThrough OpenRouter$0.15 / $4.00checked 3 hours ago1M131K max reply103 tok/sNoNoConfirmed
OpenRouterOpenRouter's own listing$1.40 / $4.40checked 26 hours ago1Mnot measuredUnknownUnknownUnknown
Mistral AInvfp4Through OpenRouter$1.40 / $4.40checked 3 hours ago1M131K max reply127 tok/sNoYes30 daysConfirmed
Together AIThrough OpenRouter$1.40 / $4.40checked 3 hours ago1M944K max reply127 tok/sNoNoConfirmed
Venice AIThrough OpenRouter$1.40 / $4.40checked 3 hours ago1M131K max reply43 tok/sNoNoConfirmed
Cloudflare Workers AIThrough OpenRouter$1.40 / $4.40checked 3 hours ago1M944K max reply33 tok/sNoYesunknown periodUnknown
AtlasCloudfp8Through OpenRouter$1.40 / $4.40checked 3 hours ago1M131K max reply50 tok/sNoYesunknown periodUnknown
Crusoefp4Through OpenRouter$1.40 / $4.40checked 3 hours ago1M944K max reply114 tok/sNoNoConfirmed
Parasailfp8Through OpenRouter$1.40 / $4.40checked 3 hours ago1M944K max reply110 tok/sNoNoConfirmed
Basetenfp4Through OpenRouter$1.40 / $4.40checked 3 hours ago1M262K max reply65 tok/sNoNoConfirmed
Baidufp8Through OpenRouter$1.40 / $4.40checked 21 hours ago1M131K max reply83 tok/sNoYesunknown periodUnknown
WaferusThrough OpenRouter$1.40 / $4.40checked 3 days ago1M944K max reply77 tok/sNoNoConfirmed
WaferThrough OpenRouter$1.82 / $4.40checked 4 days ago1M944K max reply95 tok/sNoNoConfirmed
Fireworks AIThrough OpenRouter$1.40 / $4.40checked 3 hours ago1M944K max reply43 tok/sNoNoConfirmed
Z.AIfp8Through OpenRouter$1.40 / $4.40checked 3 hours ago1M131K max reply57 tok/sNoNoConfirmed
ModalThrough OpenRouter$1.40 / $4.40checked 3 hours ago1M944K max reply62 tok/sNoNoConfirmed
PrimeIntellectThrough OpenRouter$1.40 / $4.40checked 3 hours ago1M131K max reply127 tok/sNoNoConfirmed
Basetenfast tierfp8Through OpenRouter$2.10 / $6.60checked 9 hours ago1M262K max reply141 tok/sNoNoUnknown
Alibaba Cloudfast tierThrough OpenRouter$2.80 / $8.80checked 3 hours ago1M131K max reply73 tok/sNoYesunknown periodUnknown

Across the 36 listings we hold: 35 say they do not train on prompts (2 of them only through OpenRouter), 0 say they do and 1 does not say. 26 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

API features per host
ProviderTool callingJSON outputStrict schema
RekaThrough OpenRouter✓✓✓
SiliconFlowfp8Through OpenRouter✓✓✗
InferenceNetThrough OpenRouter✓✓✓
Novita AIfp8Direct and through OpenRouter✓✓✗
DeepInfrafp4Direct and through OpenRouter✓✓✓
PhalaThrough OpenRouter✓✓✓
DigitalOcean GradientThrough OpenRouter✓✓✓
Morphfp8Through OpenRouter✓✓✓
GMICloudfp8Through OpenRouter✓✓✗
Inceptronfp4Through OpenRouter✓✓✗
AkashMLfp8Through OpenRouter✓✓✓
Decartfp4Through OpenRouter✓✓✓
Alibaba CloudThrough OpenRouter✓✓✗
Makorafp4Through OpenRouter✓✓✓
FriendliThrough OpenRouter✓✓✓
Sail Researchus · fp8Through OpenRouter✓✓✓
Sail Researchfp8Through OpenRouter✓✓✓
RelaceThrough OpenRouter✓✓✗
OpenRouterOpenRouter's own listing✓✓✓
Mistral AInvfp4Through OpenRouter✓✓✓
Together AIThrough OpenRouter✓✓✓
Venice AIThrough OpenRouter✓✓✓
Cloudflare Workers AIThrough OpenRouter✓✓✓
AtlasCloudfp8Through OpenRouter✓✓✓
Crusoefp4Through OpenRouter✓✓✓
Parasailfp8Through OpenRouter✓✓✓
Basetenfp4Through OpenRouter✓✓✓
Baidufp8Through OpenRouter✓✓✓
WaferusThrough OpenRouter✓✓✓
WaferThrough OpenRouter✓✓✓
Fireworks AIThrough OpenRouter✓✓✓
Z.AIfp8Through OpenRouter✓✓✗
ModalThrough OpenRouter✓✓✓
PrimeIntellectThrough OpenRouter✓✓✓
Basetenfast · fp8Through OpenRouter✓✗✗
Alibaba CloudfastThrough OpenRouter✓✓✗

Tool calling: 36 of 36 listings say yes. JSON output: 35 of 36 listings say yes, 1 says no. Strict schema: 27 of 36 listings say yes, 9 say no.

02

Models people weigh against GLM 5.3

03

When we formed this view

Recent changes

Oct 1, 2026Price changeHost Relace raised GLM 5.3 input and cache-read pricing by 12%
What movedinput +12% ($0.130 → $0.145 per 1M tokens), cache read +12% ($0.130 → $0.145 per 1M tokens)
Sep 30, 2026Price changeGLM 5.3 cut across 4 hosts, by up to 60% at Phala (input)
What movedGLM 5.3 moved on 4 hosts: Phala: input −60% ($1.40 → $0.56 per 1M tokens), output −23% ($4.40 → $3.40 per 1M tokens), cache read −31% ($0.26 → $0.18 per 1M tokens); Relace: input −13% ($0.15 → $0.13 per 1M tokens), cache read −13% ($0.15 → $0.13 per 1M tokens); AkashML: input −10% ($1.17 → $1.05 per 1M tokens), output −10% ($3.96 → $3.56 per 1M tokens); Makora: output −6% ($4.20 → $3.93 per 1M tokens)
Sep 29, 2026Price changeGLM 5.3 repriced across 4 hosts: Makora input down 19%, Morph input up 163%
What movedGLM 5.3 moved on 4 hosts: Morph: input +163% ($0.452 → $1.190 per 1M tokens), output +7% ($2.805 → $2.992 per 1M tokens), cache read +33% ($0.15 → $0.20 per 1M tokens); Inceptron: input +93% ($0.311 → $0.600 per 1M tokens), output +22% ($2.79 → $3.39 per 1M tokens), cache read +38% ($0.13 → $0.18 per 1M tokens); Relace: output +21% ($3.30 → $4.00 per 1M tokens), cache read +50% ($0.10 → $0.15 per 1M tokens); Makora: input −19% ($1.05 → $0.85 per 1M tokens)
Sep 28, 2026Price changeGLM 5.3 repriced across 6 hosts: Relace input down 73%, output up 94%
What movedGLM 5.3 moved on 6 hosts: Relace: input −73% ($0.55 → $0.15 per 1M tokens), output +94% ($1.70 → $3.30 per 1M tokens); Wafer: input +86% ($0.98 → $1.82 per 1M tokens); Morph: input −62% ($1.19 → $0.45 per 1M tokens), output −25% ($3.74 → $2.81 per 1M tokens), cache read −25% ($0.20 → $0.15 per 1M tokens); AtlasCloud: input −57% ($1.40 → $0.60 per 1M tokens), output −57% ($4.40 → $1.89 per 1M tokens), cache read −57% ($0.260 → $0.112 per 1M tokens); Reka: input −46% ($0.68 → $0.37 per 1M tokens), output −56% ($2.57 → $1.14 per 1M tokens), cache read −56% ($0.152 → $0.067 per 1M tokens); Inceptron: input +3% ($0.30 → $0.31 per 1M tokens), output +11% ($2.51 → $2.79 per 1M tokens), cache read +113% ($0.0623 → $0.1326 per 1M tokens)
Sep 27, 2026Price changeGLM 5.3 repriced across 3 hosts: Reka input down 48%, Inceptron output up 144%
What movedGLM 5.3 moved on 3 hosts: Inceptron: input +11% ($0.27 → $0.30 per 1M tokens), output +144% ($1.03 → $2.51 per 1M tokens), cache read +11% ($0.056 → $0.062 per 1M tokens); Reka: input −48% ($0.683 → $0.356 per 1M tokens), cache read −61% ($0.152 → $0.060 per 1M tokens); Wafer: input −30% ($1.40 → $0.98 per 1M tokens); Inceptron: input +15% ($0.238 → $0.274 per 1M tokens), output +15% ($0.898 → $1.032 per 1M tokens), cache read +15% ($0.0489 → $0.0562 per 1M tokens)
Sep 26, 2026Price changeGLM 5.3 cut across 2 hosts, by up to 43% at Inceptron (output), with cache read down 44% there
What movedGLM 5.3 moved on 2 hosts: Inceptron: input −36% ($0.597 → $0.380 per 1M tokens), output −43% ($2.50 → $1.43 per 1M tokens), cache read −44% ($0.152 → $0.085 per 1M tokens); Relace: input −21% ($0.70 → $0.55 per 1M tokens), output −23% ($2.20 → $1.70 per 1M tokens), cache read −23% ($0.13 → $0.10 per 1M tokens)
Sep 25, 2026Price changeGLM 5.3 repriced across 3 hosts: Sail Research (US region) input down 36%, output up 3%
What movedGLM 5.3 moved on 3 hosts: Sail Research (US region): input −36% ($1.21 → $0.77 per 1M tokens), output +3% ($3.87 → $4.00 per 1M tokens), cache read −16% ($0.225 → $0.190 per 1M tokens); Inceptron: input −24% ($0.79 → $0.60 per 1M tokens), output −5% ($2.647 → $2.502 per 1M tokens), cache read −25% ($0.20 → $0.15 per 1M tokens); Io Net: input −1% ($0.78 → $0.77 per 1M tokens)
Sep 25, 2026BenchmarkScored 0.029 via Max on Arena Agent
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.008 via Max on Arena Agent · Recovery
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.03 via Max on Arena Agent · Steerability
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 36 listings does not say whether it trains on prompts, and 2 answer only through OpenRouter, not for their own listing.
  • We hold no batch or off-peak rate for any of its listings.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Text in, text out
Catalogue slug
z-ai-glm-5-3

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us