GLM 5.3
Z.ai · released Aug 18, 2026
- Type
- Closed
- Input
- $1.40
- Output
- $4.40
- Cached
- $0.26
List price · per 1M tokens · Z.AI at 1M context · source ↗ · a reseller below undercuts it; the table carries the spread
Our take
Written Sep 30, 2026GLM 5.3 is a hosted-only text model: we list no download for it, so using it means choosing a host. It is strongest where people rate the answers themselves, and mid-field on the harder task sets.
Reach for it on chat, writing and general assistant work where human-rated answer quality matters, or in agent sessions that need reliable tool selection, where it places 9th of 55 on Arena Agent · Tool use via Max as of 25 Sep 2026. Long documents need not be split up first, though recall across all of it is unverified in our data. Skip it if you need to run the model on your own hardware, or if data analysis is the job.
The case for it
- 16th of 168 on Arena Text (overall) via Max as of 25 Sep 2026, a board that records which answer people preferred rather than whether it was correct.
- Among the better measured models at picking the right tool in an agent session: 9th of 55 on Arena Agent · Tool use via Max as of 25 Sep 2026, a board that scores calling the right tool and not inventing one.
- 87.9% on LiveBench Mathematics and 85.8% on LiveBench Reasoning, both averages over competition-style and monthly-refreshed task sets rather than work in an existing project.
- The request capacity takes a long report or a stack of documents alongside the question, so long inputs need not be split up first.
The case against it
- We list no download for it, so every route we hold is a hosted offer, and the licence is not disclosed.
- 44th of 58 on LiveBench Data Analysis as of 25 Jun 2026, a board of table and event-ordering tasks, so spreadsheet-style work is where it trails its own other results.
- 60.91% on LiveBench Agentic Coding against 78.95% on LiveBench Coding, and the agentic figure was run inside an agent harness, so it reflects the model in that scaffold rather than on its own.
How good is it?
A closed text model for everyday questions, drafting, coding and calling tools.
- getting answers to everyday questionsArena Text (overall) · 16th of 168
- drafts, rewrites and editingArena Creative Writing · 19th of 168
- writing and completing codeArena Coding · 23rd of 168
- calling tools to carry out requestsArena Agent · Tool use · 9th of 55
EverydayGeneral questions and everyday reasoning
Arena Text (overall)16th of 168 · 1480
CodingWriting and fixing code on its own
Arena Coding23rd of 168 · 1522
AgenticPlanning, calling tools, staying on task
Arena Agent17th of 55 · 0.029
WritingDrafting and rewriting prose
Arena Creative Writing19th of 168 · 1456
Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.
Boards this model appears on that none of the ratings above are built on.
Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Where to rent it
Prices checked between 3 hours and 4 days ago — each listing carries its own date.
Some hosts sell this model at two prices: on their own price list (“direct”) and on their OpenRouter listing (“through OpenRouter”). Where the two differ, the row shows both, each with the date we last read it.
InferenceNet, through OpenRouter
Why this differs from the header. The strip above quotes Z.AI's own list price; this is the cheapest live offer, whoever is serving it — a reseller undercutting a lab is ordinary commerce, not an error.
- per 1M tokens
- $0.68 in / $2.28 out
- Context served
- 1M
- Throughput
- ~91 tok/s
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| RekaThrough OpenRouter | $0.37 / $1.14checked 9 hours ago | 262K236K max reply | 60 tok/s | No | No | Confirmed |
| SiliconFlowfp8Through OpenRouter | $0.70 / $2.20checked 3 hours ago | 1M262K max reply | 54 tok/s | No | No | Confirmed |
| InferenceNetThrough OpenRouter | $0.68 / $2.28checked 2 days ago | 1M131K max reply | 91 tok/s | No | No | Confirmed |
| Novita AIfp8Direct and through OpenRouter | $1.40 / $4.40directchecked 3 hours ago$0.78 / $2.46through OpenRouterchecked 3 hours ago | 1M131K max reply through OpenRouter | 45 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| DeepInfrafp4Direct and through OpenRouter | $0.90 / $4.00directchecked 3 hours ago$0.56 / $2.50through OpenRouterchecked 3 hours ago | 1M131K max reply through OpenRouter | 36 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| PhalaThrough OpenRouter | $0.84 / $2.64checked 3 hours ago | 1M944K max reply | 44 tok/s | No | No | Confirmed |
| DigitalOcean GradientThrough OpenRouter | $0.91 / $2.86checked 3 hours ago | 1M128K max reply | 47 tok/s | No | No | Confirmed |
| Morphfp8Through OpenRouter | $1.19 / $2.99checked 2 days ago | 1M944K max reply | 109 tok/s | No | No | Confirmed |
| GMICloudfp8Through OpenRouter | $0.98 / $3.08checked 3 hours ago | 1M944K max reply | 53 tok/s | No | Yesunknown period | Unknown |
| Inceptronfp4Through OpenRouter | $0.60 / $3.39checked 3 hours ago | 1M944K max reply | 60 tok/s | No | No | Confirmed |
| AkashMLfp8Through OpenRouter | $1.05 / $3.56checked 3 hours ago | 1M131K max reply | 80 tok/s | No | No | Confirmed |
| Decartfp4Through OpenRouter | $1.19 / $3.74checked 3 hours ago | 1M944K max reply | 129 tok/s | No | No | Confirmed |
| Alibaba CloudThrough OpenRouter | $1.19 / $3.74checked 3 hours ago | 1M131K max reply | 74 tok/s | No | Yesunknown period | Unknown |
| Makorafp4Through OpenRouter | $0.85 / $3.93checked 9 hours ago | 980K128K max reply | 109 tok/s | No | No | Confirmed |
| FriendliThrough OpenRouter | $1.26 / $3.96checked 3 hours ago | 1M944K max reply | 92 tok/s | No | Yesunknown period | Unknown |
| Sail Researchusfp8Through OpenRouter | $0.77 / $4.00checked 3 days ago | 1M944K max reply | 83 tok/s | No | No | Unknown |
| Sail Researchfp8Through OpenRouter | $0.77 / $4.00checked 3 days ago | 1M944K max reply | 74 tok/s | No | No | Confirmed |
| RelaceThrough OpenRouter | $0.15 / $4.00checked 3 hours ago | 1M131K max reply | 103 tok/s | No | No | Confirmed |
| OpenRouterOpenRouter's own listing | $1.40 / $4.40checked 26 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| Mistral AInvfp4Through OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M131K max reply | 127 tok/s | No | Yes30 days | Confirmed |
| Together AIThrough OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M944K max reply | 127 tok/s | No | No | Confirmed |
| Venice AIThrough OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M131K max reply | 43 tok/s | No | No | Confirmed |
| Cloudflare Workers AIThrough OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M944K max reply | 33 tok/s | No | Yesunknown period | Unknown |
| AtlasCloudfp8Through OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M131K max reply | 50 tok/s | No | Yesunknown period | Unknown |
| Crusoefp4Through OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M944K max reply | 114 tok/s | No | No | Confirmed |
| Parasailfp8Through OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M944K max reply | 110 tok/s | No | No | Confirmed |
| Basetenfp4Through OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M262K max reply | 65 tok/s | No | No | Confirmed |
| Baidufp8Through OpenRouter | $1.40 / $4.40checked 21 hours ago | 1M131K max reply | 83 tok/s | No | Yesunknown period | Unknown |
| WaferusThrough OpenRouter | $1.40 / $4.40checked 3 days ago | 1M944K max reply | 77 tok/s | No | No | Confirmed |
| WaferThrough OpenRouter | $1.82 / $4.40checked 4 days ago | 1M944K max reply | 95 tok/s | No | No | Confirmed |
| Fireworks AIThrough OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M944K max reply | 43 tok/s | No | No | Confirmed |
| Z.AIfp8Through OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M131K max reply | 57 tok/s | No | No | Confirmed |
| ModalThrough OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M944K max reply | 62 tok/s | No | No | Confirmed |
| PrimeIntellectThrough OpenRouter | $1.40 / $4.40checked 3 hours ago | 1M131K max reply | 127 tok/s | No | No | Confirmed |
| Basetenfast tierfp8Through OpenRouter | $2.10 / $6.60checked 9 hours ago | 1M262K max reply | 141 tok/s | No | No | Unknown |
| Alibaba Cloudfast tierThrough OpenRouter | $2.80 / $8.80checked 3 hours ago | 1M131K max reply | 73 tok/s | No | Yesunknown period | Unknown |
Across the 36 listings we hold: 35 say they do not train on prompts (2 of them only through OpenRouter), 0 say they do and 1 does not say. 26 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| RekaThrough OpenRouter | ✓ | ✓ | ✓ |
| SiliconFlowfp8Through OpenRouter | ✓ | ✓ | ✗ |
| InferenceNetThrough OpenRouter | ✓ | ✓ | ✓ |
| Novita AIfp8Direct and through OpenRouter | ✓ | ✓ | ✗ |
| DeepInfrafp4Direct and through OpenRouter | ✓ | ✓ | ✓ |
| PhalaThrough OpenRouter | ✓ | ✓ | ✓ |
| DigitalOcean GradientThrough OpenRouter | ✓ | ✓ | ✓ |
| Morphfp8Through OpenRouter | ✓ | ✓ | ✓ |
| GMICloudfp8Through OpenRouter | ✓ | ✓ | ✗ |
| Inceptronfp4Through OpenRouter | ✓ | ✓ | ✗ |
| AkashMLfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Decartfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Alibaba CloudThrough OpenRouter | ✓ | ✓ | ✗ |
| Makorafp4Through OpenRouter | ✓ | ✓ | ✓ |
| FriendliThrough OpenRouter | ✓ | ✓ | ✓ |
| Sail Researchus · fp8Through OpenRouter | ✓ | ✓ | ✓ |
| Sail Researchfp8Through OpenRouter | ✓ | ✓ | ✓ |
| RelaceThrough OpenRouter | ✓ | ✓ | ✗ |
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| Mistral AInvfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Together AIThrough OpenRouter | ✓ | ✓ | ✓ |
| Venice AIThrough OpenRouter | ✓ | ✓ | ✓ |
| Cloudflare Workers AIThrough OpenRouter | ✓ | ✓ | ✓ |
| AtlasCloudfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Crusoefp4Through OpenRouter | ✓ | ✓ | ✓ |
| Parasailfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Basetenfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Baidufp8Through OpenRouter | ✓ | ✓ | ✓ |
| WaferusThrough OpenRouter | ✓ | ✓ | ✓ |
| WaferThrough OpenRouter | ✓ | ✓ | ✓ |
| Fireworks AIThrough OpenRouter | ✓ | ✓ | ✓ |
| Z.AIfp8Through OpenRouter | ✓ | ✓ | ✗ |
| ModalThrough OpenRouter | ✓ | ✓ | ✓ |
| PrimeIntellectThrough OpenRouter | ✓ | ✓ | ✓ |
| Basetenfast · fp8Through OpenRouter | ✓ | ✗ | ✗ |
| Alibaba CloudfastThrough OpenRouter | ✓ | ✓ | ✗ |
Tool calling: 36 of 36 listings say yes. JSON output: 35 of 36 listings say yes, 1 says no. Strict schema: 27 of 36 listings say yes, 9 say no.
Models people weigh against GLM 5.3
When we formed this view
Recent changes
What moved
input +12% ($0.130 → $0.145 per 1M tokens), cache read +12% ($0.130 → $0.145 per 1M tokens)What moved
GLM 5.3 moved on 4 hosts: Phala: input −60% ($1.40 → $0.56 per 1M tokens), output −23% ($4.40 → $3.40 per 1M tokens), cache read −31% ($0.26 → $0.18 per 1M tokens); Relace: input −13% ($0.15 → $0.13 per 1M tokens), cache read −13% ($0.15 → $0.13 per 1M tokens); AkashML: input −10% ($1.17 → $1.05 per 1M tokens), output −10% ($3.96 → $3.56 per 1M tokens); Makora: output −6% ($4.20 → $3.93 per 1M tokens)What moved
GLM 5.3 moved on 4 hosts: Morph: input +163% ($0.452 → $1.190 per 1M tokens), output +7% ($2.805 → $2.992 per 1M tokens), cache read +33% ($0.15 → $0.20 per 1M tokens); Inceptron: input +93% ($0.311 → $0.600 per 1M tokens), output +22% ($2.79 → $3.39 per 1M tokens), cache read +38% ($0.13 → $0.18 per 1M tokens); Relace: output +21% ($3.30 → $4.00 per 1M tokens), cache read +50% ($0.10 → $0.15 per 1M tokens); Makora: input −19% ($1.05 → $0.85 per 1M tokens)What moved
GLM 5.3 moved on 6 hosts: Relace: input −73% ($0.55 → $0.15 per 1M tokens), output +94% ($1.70 → $3.30 per 1M tokens); Wafer: input +86% ($0.98 → $1.82 per 1M tokens); Morph: input −62% ($1.19 → $0.45 per 1M tokens), output −25% ($3.74 → $2.81 per 1M tokens), cache read −25% ($0.20 → $0.15 per 1M tokens); AtlasCloud: input −57% ($1.40 → $0.60 per 1M tokens), output −57% ($4.40 → $1.89 per 1M tokens), cache read −57% ($0.260 → $0.112 per 1M tokens); Reka: input −46% ($0.68 → $0.37 per 1M tokens), output −56% ($2.57 → $1.14 per 1M tokens), cache read −56% ($0.152 → $0.067 per 1M tokens); Inceptron: input +3% ($0.30 → $0.31 per 1M tokens), output +11% ($2.51 → $2.79 per 1M tokens), cache read +113% ($0.0623 → $0.1326 per 1M tokens)What moved
GLM 5.3 moved on 3 hosts: Inceptron: input +11% ($0.27 → $0.30 per 1M tokens), output +144% ($1.03 → $2.51 per 1M tokens), cache read +11% ($0.056 → $0.062 per 1M tokens); Reka: input −48% ($0.683 → $0.356 per 1M tokens), cache read −61% ($0.152 → $0.060 per 1M tokens); Wafer: input −30% ($1.40 → $0.98 per 1M tokens); Inceptron: input +15% ($0.238 → $0.274 per 1M tokens), output +15% ($0.898 → $1.032 per 1M tokens), cache read +15% ($0.0489 → $0.0562 per 1M tokens)What moved
GLM 5.3 moved on 2 hosts: Inceptron: input −36% ($0.597 → $0.380 per 1M tokens), output −43% ($2.50 → $1.43 per 1M tokens), cache read −44% ($0.152 → $0.085 per 1M tokens); Relace: input −21% ($0.70 → $0.55 per 1M tokens), output −23% ($2.20 → $1.70 per 1M tokens), cache read −23% ($0.13 → $0.10 per 1M tokens)What moved
GLM 5.3 moved on 3 hosts: Sail Research (US region): input −36% ($1.21 → $0.77 per 1M tokens), output +3% ($3.87 → $4.00 per 1M tokens), cache read −16% ($0.225 → $0.190 per 1M tokens); Inceptron: input −24% ($0.79 → $0.60 per 1M tokens), output −5% ($2.647 → $2.502 per 1M tokens), cache read −25% ($0.20 → $0.15 per 1M tokens); Io Net: input −1% ($0.78 → $0.77 per 1M tokens)Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 1 of 36 listings does not say whether it trains on prompts, and 2 answer only through OpenRouter, not for their own listing.
- We hold no batch or off-peak rate for any of its listings.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.
Identifiers
- Takes in, gives back
- Text in, text out
- Catalogue slug
- z-ai-glm-5-3