GLM 5.3 Flash cut across 3 hosts, by up to 50% at OpenInference (input)
Published 25 September 2026
Price move, per 1M tokens
- OpenInference · input↓ 50%$0.100 → $0.050
- OpenInference · outputunchanged$0.50 → $0.50
- OpenInference · cache read↓ 40%$0.025 → $0.015
- Inceptron · input↓ 27%$0.15 → $0.11
- Inceptron · output↓ 10%$0.50 → $0.45
- Inceptron · cache readunchanged$0.070 → $0.070
- Wafer · input↓ 22%$0.089 → $0.069
- Wafer · outputunchanged$0.35 → $0.35
- Wafer · cache readunchanged$0.030 → $0.030
| host | rate | was | now | change |
|---|---|---|---|---|
| OpenInference | input | $0.100 | $0.050 | ↓ 50% |
| OpenInference | output | $0.50 | $0.50 | — |
| OpenInference | cache read | $0.025 | $0.015 | ↓ 40% |
| Inceptron | input | $0.15 | $0.11 | ↓ 27% |
| Inceptron | output | $0.50 | $0.45 | ↓ 10% |
| Inceptron | cache read | $0.070 | $0.070 | — |
| Wafer | input | $0.089 | $0.069 | ↓ 22% |
| Wafer | output | $0.35 | $0.35 | — |
| Wafer | cache read | $0.030 | $0.030 | — |
OpenInference input: ↓ 50%Inceptron input: ↓ 27%Wafer input: ↓ 22%
If you buy
Renting this model at Wafer now costs less to run, so heavier prompt and cache traffic fits a tighter budget; re-check your usage tier and cache settings before your next top-up.
The model
GLM 5.3 Flashopen the model →
Open weights321.3B parameters1311K context35 hosts
intelligence26th of 168 writing43rd of 168 coding22nd of 168 agents27th of 55
GLM 5.3 Flash is a downloadable model you can run yourself, and it is at its best when the job is agentic: picking the right tool and finishing the task. It is at its worst when the job is holding to a format under instruction, and we list no licence for it, so the terms need checking at the source.
Source
Read on openrouter.ai · we published this on 25 September 2026.
read the original at openrouter.ai ↗