Provider guide
Where to run GLM 5.3 Flash
33 live listings tracked — output prices vary 4.0× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.
DeepInfrazero-retention through OpenRouterDirect and through OpenRouter$0.15 in/1M$0.075 in/1M$0.50 out/1Mdirect$0.25 out/1Mthrough OpenRouter18 tok/sthrough OpenRouter1Mfp44 hours agoNovita AIzero-retention through OpenRouterDirect and through OpenRouter$0.15 in/1M$0.084 in/1M$0.50 out/1Mdirect$0.28 out/1Mthrough OpenRouter17 tok/sthrough OpenRouter1Mfp84 hours agoStreamLake$0.087 in/1M$0.29 out/1M36 tok/s1Mfp84 hours agoOpenInferencezero-retention$0.020 in/1M$0.30 out/1M15 tok/s1Mfp44 hours agoGMICloud$0.090 in/1M$0.30 out/1M29 tok/s1Mfp84 hours agoNear AIzero-retention$0.10 in/1M$0.35 out/1M10 tok/s1Mfp84 hours agoPhalazero-retention$0.12 in/1M$0.40 out/1M23 tok/s1Mfp84 hours agoDecartzero-retention$0.13 in/1M$0.42 out/1M73 tok/s1Mfp44 hours agoIo Netzero-retention$0.14 in/1M$0.45 out/1M31 tok/s262Kfp84 hours agoInceptronzero-retention$0.23 in/1M$0.45 out/1M15 tok/s1Mfp84 hours agoModalzero-retention$0.15 in/1M$0.50 out/1M84 tok/s1Mnvfp44 hours agoRekazero-retention$0.15 in/1M$0.50 out/1M54 tok/s262K—4 hours agoRelacezero-retentionCHEAPEST$0.035 in/1M$0.50 out/1M40 tok/s1M—4 hours agoSiliconFlowzero-retention$0.15 in/1M$0.50 out/1M35 tok/s1Mfp84 hours agoCrusoezero-retention$0.15 in/1M$0.50 out/1M81 tok/s1Mfp44 hours agoDigitalOcean Gradientzero-retention$0.15 in/1M$0.50 out/1M24 tok/s1M—4 hours agoParasailzero-retention$0.15 in/1M$0.50 out/1M65 tok/s1Mfp84 hours agoTogether AIzero-retention$0.15 in/1M$0.50 out/1M81 tok/s1M—4 hours agoOpenRouter$0.15 in/1M$0.50 out/1M—1M—4 hours agoCoreWeavezero-retention$0.15 in/1M$0.50 out/1M82 tok/s1Mnvfp44 hours agoBasetenzero-retention$0.15 in/1M$0.50 out/1M100 tok/s1Mfp84 hours agoAtlasCloud$0.15 in/1M$0.50 out/1M34 tok/s1Mfp84 hours agoVenice AIzero-retention$0.15 in/1M$0.50 out/1M21 tok/s1M—4 hours agoFireworks AIzero-retention$0.15 in/1M$0.50 out/1M48 tok/s1M—4 hours agoFriendli$0.15 in/1M$0.50 out/1M77 tok/s1M—4 hours agoZ.AIzero-retention$0.15 in/1M$0.50 out/1M36 tok/s1Mfp84 hours agoNextBitzero-retention$0.17 in/1M$0.55 out/1M31 tok/s1Mfp84 hours agoInferenceNetzero-retention$0.050 in/1M$0.60 out/1M34 tok/s1Mfp44 hours agoSail Research$0.045 in/1M$0.60 out/1M31 tok/s1Mfp44 hours agoMorphzero-retention$0.20 in/1M$0.70 out/1M106 tok/s1Mfp84 hours agoFireworks AIzero-retention$0.23 in/1M$0.75 out/1M70 tok/s1M—4 hours agoWaferzero-retention$1.00 in/1M$0.75 out/1M30 tok/s1M—4 hours agoCloudflare Workers AI$0.30 in/1M$1.00 out/1M28 tok/s1M—4 hours ago
Full specs, hardware verdicts and benchmarks on the GLM 5.3 Flash model page →