GLM 5.3 Flash repriced across 2 hosts: InferenceNet input down 50%, Wafer output up 50%
Published 1 October 2026
Price move, per 1M tokens
- Wafer · input↑ 33%$0.75 → $1.00
- Wafer · output↑ 50%$0.50 → $0.75
- Wafer · cache readunchanged$0.030 → $0.030
- InferenceNet · input↓ 50%$0.100 → $0.050
- InferenceNet · output↑ 33%$0.45 → $0.60
- InferenceNet · cache read↑ 60%$0.030 → $0.048
| host | rate | was | now | change |
|---|---|---|---|---|
| Wafer | input | $0.75 | $1.00 | ↑ 33% |
| Wafer | output | $0.50 | $0.75 | ↑ 50% |
| Wafer | cache read | $0.030 | $0.030 | — |
| InferenceNet | input | $0.100 | $0.050 | ↓ 50% |
| InferenceNet | output | $0.45 | $0.60 | ↑ 33% |
| InferenceNet | cache read | $0.030 | $0.048 | ↑ 60% |
Wafer output: ↑ 50%InferenceNet input: ↓ 50%InferenceNet output: ↑ 33%
If you buy
On Wafer, this model now costs more to run, so check whether your workload leans on output before your next top-up.
The model
GLM 5.3 Flashopen the model →
Open weights321.3B parameters1311K context35 hosts
intelligence26th of 168 writing43rd of 168 coding22nd of 168 agents27th of 55
GLM 5.3 Flash is a downloadable model you can run yourself, and it is at its best when the job is agentic: picking the right tool and finishing the task. It is at its worst when the job is holding to a format under instruction, and we list no licence for it, so the terms need checking at the source.
Source
Read on openrouter.ai · we published this on 1 October 2026.
read the original at openrouter.ai ↗