GLM 5.3 Flash repriced across 3 hosts: Relace input down 43%, Wafer input up 523%
Published 26 September 2026
Price move, per 1M tokens
- Wafer · input↑ 523%$0.069 → $0.430
- Wafer · output↑ 43%$0.35 → $0.50
- Wafer · cache readunchanged$0.030 → $0.030
- Relace · input↓ 43%$0.070 → $0.040
- Relace · output↑ 79%$0.28 → $0.50
- Relace · cache read↓ 25%$0.020 → $0.015
- OpenInference · inputunchanged$0.050 → $0.050
- OpenInference · outputunchanged$0.50 → $0.50
- OpenInference · cache read↑ 33%$0.015 → $0.020
| host | rate | was | now | change |
|---|---|---|---|---|
| Wafer | input | $0.069 | $0.430 | ↑ 523% |
| Wafer | output | $0.35 | $0.50 | ↑ 43% |
| Wafer | cache read | $0.030 | $0.030 | — |
| Relace | input | $0.070 | $0.040 | ↓ 43% |
| Relace | output | $0.28 | $0.50 | ↑ 79% |
| Relace | cache read | $0.020 | $0.015 | ↓ 25% |
| OpenInference | input | $0.050 | $0.050 | — |
| OpenInference | output | $0.50 | $0.50 | — |
| OpenInference | cache read | $0.015 | $0.020 | ↑ 33% |
Wafer input: ↑ 523%Relace input: ↓ 43%Relace output: ↑ 79%OpenInference cache read: ↑ 33%
If you buy
Renting this model on Wafer now costs far more per token, so re-check your usage and whether a different host fits your workload before your next top-up.
The model
GLM 5.3 Flashopen the model →
Open weights321.3B parameters1311K context35 hosts
intelligence26th of 168 writing43rd of 168 coding22nd of 168 agents27th of 55
GLM 5.3 Flash is a downloadable model you can run yourself, and it is at its best when the job is agentic: picking the right tool and finishing the task. It is at its worst when the job is holding to a format under instruction, and we list no licence for it, so the terms need checking at the source.
Source
Read from the hosts' own published rates · we published this on 26 September 2026.