News / PRICE

GLM 5.3 Flash repriced across 5 hosts: OpenInference output down 18%, InferenceNet output up 61%

Published 30 September 2026
Price move, per 1M tokens
  • InferenceNet · input↑ 11%
    $0.090 → $0.100
  • InferenceNet · output↑ 61%
    $0.28 → $0.45
  • InferenceNet · cache read↑ 50%
    $0.020 → $0.030
  • Inceptron · input↑ 50%
    $0.150 → $0.225
  • Inceptron · outputunchanged
    $0.45 → $0.45
  • Inceptron · cache read↑ 33%
    $0.060 → $0.080
  • Relace · input↑ 40%
    $0.025 → $0.035
  • Relace · outputunchanged
    $0.50 → $0.50
  • Relace · cache read↑ 40%
    $0.025 → $0.035
  • OpenInference · inputunchanged
    $0.020 → $0.020
  • OpenInference · output↓ 18%
    $0.300 → $0.247
  • OpenInference · cache readunchanged
    $0.010 → $0.010
  • Morph · input↓ 10%
    $0.20 → $0.18
  • Morph · output↓ 10%
    $0.70 → $0.63
  • Morph · cache readunchanged
    $0.040 → $0.040
InferenceNet output: ↑ 61%Inceptron input: ↑ 50%Relace input and cache read: ↑ 40%OpenInference output: ↓ 18%Morph input and output: ↓ 10%
If you buy

Output-heavy workloads on this model here now cost more per call, so re-check your spend cap and whether shorter replies or another host fit your use.

The model
GLM 5.3 Flashopen the model →
Open weights321.3B parameters1311K context35 hosts
intelligence26th of 168 writing43rd of 168 coding22nd of 168 agents27th of 55

GLM 5.3 Flash is a downloadable model you can run yourself, and it is at its best when the job is agentic: picking the right tool and finishing the task. It is at its worst when the job is holding to a format under instruction, and we list no licence for it, so the terms need checking at the source.

Source

Read on openrouter.ai · we published this on 30 September 2026.

read the original at openrouter.ai ↗