News / PRICE

GLM 5.3 Flash repriced across 4 hosts: OpenInference input down 60%, Inceptron input up 25%

Published 29 September 2026
Price move, per 1M tokens
  • OpenInference · input↓ 60%
    $0.050 → $0.020
  • OpenInference · output↓ 40%
    $0.50 → $0.30
  • OpenInference · cache read↓ 50%
    $0.020 → $0.010
  • Relace · input↓ 38%
    $0.040 → $0.025
  • Relace · outputunchanged
    $0.50 → $0.50
  • Relace · cache read↑ 67%
    $0.015 → $0.025
  • Morph · inputunchanged
    $0.15 → $0.15
  • Morph · outputunchanged
    $0.54 → $0.54
  • Morph · cache read↑ 31%
    $0.0306 → $0.0400
  • Inceptron · input↑ 25%
    $0.12 → $0.15
  • Inceptron · outputunchanged
    $0.45 → $0.45
  • Inceptron · cache read↑ 50%
    $0.040 → $0.060
OpenInference input: ↓ 60%Relace input: ↓ 38%Relace cache read: ↑ 67%Morph cache read: ↑ 31%Inceptron input: ↑ 25%
If you buy

If you run this model on Morph, re-check your cache-read spend before your next top-up, since prompt reuse now costs more there.

The model
GLM 5.3 Flashopen the model →
Open weights321.3B parameters1311K context35 hosts
intelligence26th of 168 writing43rd of 168 coding22nd of 168 agents27th of 55

GLM 5.3 Flash is a downloadable model you can run yourself, and it is at its best when the job is agentic: picking the right tool and finishing the task. It is at its worst when the job is holding to a format under instruction, and we list no licence for it, so the terms need checking at the source.

Source

Read on openrouter.ai · we published this on 29 September 2026.

read the original at openrouter.ai ↗