Models /Llama 3.1 8B Instruct /Where to run
Provider guide

Where to run Llama 3.1 8B Instruct

7 live offers tracked — output prices vary 5.7× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.

CHEAPEST
$0.050 / $0.080
per 1M tokens in / out
FASTEST MEASURED
306 tok/s
measured throughput · $0.050 input / $0.080 output per million tokens
Novita AI$0.020 in/1M$0.050 out/1M16K9h agoOpenRouterCHEAPEST$0.050 in/1M$0.080 out/1M131K3h agoCloudflare Workers AI$0.15 in/1M$0.29 out/1M19 tok/s32Kfp83h agoDeepInfrazero-retention$0.020 in/1M$0.040 out/1M16 tok/s131Kfp83h agoNovita AIzero-retention$0.020 in/1M$0.050 out/1M54 tok/s16Kfp83h agoGroqzero-retention$0.050 in/1M$0.080 out/1M306 tok/s131K3h agoCoreWeavezero-retention$0.22 in/1M$0.22 out/1M119 tok/s128Kbf163h ago
Full specs, hardware verdicts and benchmarks on theLlama 3.1 8B Instruct model page →