Provider guide
Where to run GLM 4.7
10 live offers tracked — output prices vary 1.9× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.
FASTEST MEASURED
225 tok/s
measured throughput · $2.25 input / $2.75 output per million tokens
OpenRouterCHEAPEST$0.40 in/1M$1.75 out/1M—205K—41m agoDeepInfra$0.40 in/1M$1.75 out/1M—203Kfp47h agoAtlasCloud$0.52 in/1M$1.85 out/1M37 tok/s203Kfp87h agoNovita AI$0.60 in/1M$2.20 out/1M—205K—7h agoPhala$0.85 in/1M$3.30 out/1M33 tok/s131K—6d agoDeepInfrazero-retention$0.40 in/1M$1.75 out/1M30 tok/s203Kfp439m agoNovita AIzero-retention$0.54 in/1M$1.98 out/1M28 tok/s205Kfp87h agoGoogle Vertex AIzero-retention$0.60 in/1M$2.20 out/1M111 tok/s200K—7h agoVenice AIzero-retention$0.55 in/1M$2.65 out/1M26 tok/s198Kfp439m agoCerebraszero-retention$2.25 in/1M$2.75 out/1M225 tok/s131Kfp1639m ago
Full specs, hardware verdicts and benchmarks on theGLM 4.7 model page →