Models /Gemma 4 31B /Where to run
Provider guide

Where to run Gemma 4 31B

22 live offers tracked — output prices vary 3.4× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.

CHEAPEST
$0.10 / $0.34
per 1M tokens in / out
FASTEST MEASURED
121 tok/s
measured throughput · $0.38 input / $1.15 output per million tokens
OpenRouterCHEAPEST$0.10 in/1M$0.34 out/1M262K52m agoDeepInfra$0.090 in/1M$0.34 out/1M262Kfp46d agoChutes$0.12 in/1M$0.37 out/1M14 tok/s131Kfp415h agoDeepInfra$0.13 in/1M$0.38 out/1M262Kfp87h agoNovita AI$0.14 in/1M$0.40 out/1M262K7h agoFriendli$0.14 in/1M$0.40 out/1M111 tok/s262K52m agoSambaNova$0.38 in/1M$1.15 out/1M131K7h agoCoreWeavezero-retention$0.10 in/1M$0.34 out/1M39 tok/s262Kbf167h agoDeepInfraturbo tier$0.090 in/1M$0.34 out/1M50 tok/s262Kfp452m agoOpenInferencezero-retention$0.10 in/1M$0.35 out/1M54 tok/s262Kbf1652m agoVenice AIzero-retention$0.12 in/1M$0.36 out/1M38 tok/s256Kbf167h agoDeepInfrazero-retention$0.13 in/1M$0.38 out/1M49 tok/s262Kfp852m agoNovita AIzero-retention$0.14 in/1M$0.40 out/1M16 tok/s262Kbf1652m agoMorphzero-retention$0.14 in/1M$0.40 out/1M16 tok/s175Kfp47h agoSiliconFlowzero-retention$0.13 in/1M$0.40 out/1M26 tok/s262Kfp87h agoCrusoezero-retention$0.14 in/1M$0.40 out/1M34 tok/s262K7h agoParasailzero-retention$0.15 in/1M$0.40 out/1M30 tok/s262Kfp852m agoPhalazero-retention$0.15 in/1M$0.46 out/1M21 tok/s262K52m agoModelRunzero-retention$0.22 in/1M$0.55 out/1M61 tok/s262Kfp452m agoTogether AIzero-retention$0.28 in/1M$0.86 out/1M18 tok/s262K9h agoSambaNovazero-retention$0.38 in/1M$1.15 out/1M121 tok/s131K52m agoCerebraszero-retention$0.99 in/1M$1.49 out/1M29 tok/s131Kfp1652m ago
Full specs, hardware verdicts and benchmarks on theGemma 4 31B model page →