GLM 5.3 Flash repriced across 5 hosts: OpenInference output down 18%, InferenceNet output up 61%
Published 30 September 2026
Price move, per 1M tokens
- InferenceNet · input↑ 11%$0.090 → $0.100
- InferenceNet · output↑ 61%$0.28 → $0.45
- InferenceNet · cache read↑ 50%$0.020 → $0.030
- Inceptron · input↑ 50%$0.150 → $0.225
- Inceptron · outputunchanged$0.45 → $0.45
- Inceptron · cache read↑ 33%$0.060 → $0.080
- Relace · input↑ 40%$0.025 → $0.035
- Relace · outputunchanged$0.50 → $0.50
- Relace · cache read↑ 40%$0.025 → $0.035
- OpenInference · inputunchanged$0.020 → $0.020
- OpenInference · output↓ 18%$0.300 → $0.247
- OpenInference · cache readunchanged$0.010 → $0.010
- Morph · input↓ 10%$0.20 → $0.18
- Morph · output↓ 10%$0.70 → $0.63
- Morph · cache readunchanged$0.040 → $0.040
| host | rate | was | now | change |
|---|---|---|---|---|
| InferenceNet | input | $0.090 | $0.100 | ↑ 11% |
| InferenceNet | output | $0.28 | $0.45 | ↑ 61% |
| InferenceNet | cache read | $0.020 | $0.030 | ↑ 50% |
| Inceptron | input | $0.150 | $0.225 | ↑ 50% |
| Inceptron | output | $0.45 | $0.45 | — |
| Inceptron | cache read | $0.060 | $0.080 | ↑ 33% |
| Relace | input | $0.025 | $0.035 | ↑ 40% |
| Relace | output | $0.50 | $0.50 | — |
| Relace | cache read | $0.025 | $0.035 | ↑ 40% |
| OpenInference | input | $0.020 | $0.020 | — |
| OpenInference | output | $0.300 | $0.247 | ↓ 18% |
| OpenInference | cache read | $0.010 | $0.010 | — |
| Morph | input | $0.20 | $0.18 | ↓ 10% |
| Morph | output | $0.70 | $0.63 | ↓ 10% |
| Morph | cache read | $0.040 | $0.040 | — |
InferenceNet output: ↑ 61%Inceptron input: ↑ 50%Relace input and cache read: ↑ 40%OpenInference output: ↓ 18%Morph input and output: ↓ 10%
If you buy
Output-heavy workloads on this model here now cost more per call, so re-check your spend cap and whether shorter replies or another host fit your use.
The model
GLM 5.3 Flashopen the model →
Open weights321.3B parameters1311K context35 hosts
intelligence26th of 168 writing43rd of 168 coding22nd of 168 agents27th of 55
GLM 5.3 Flash is a downloadable model you can run yourself, and it is at its best when the job is agentic: picking the right tool and finishing the task. It is at its worst when the job is holding to a format under instruction, and we list no licence for it, so the terms need checking at the source.
Source
Read on openrouter.ai · we published this on 30 September 2026.
read the original at openrouter.ai ↗