Models /Llama 3.1 8B Instruct /Where to run
Provider guide

Where to run Llama 3.1 8B Instruct

6 live listings tracked — output prices vary 7.2× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.

CHEAPEST
$0.050 / $0.080
per 1M tokens in / out
FASTEST MEASURED
124 tok/s
measured throughput · $0.050 input / $0.080 output per million tokens
DeepInfrazero-retention$0.020 in/1M$0.040 out/1M28 tok/s131Kfp829 min agoNovita AIzero-retention through OpenRouterDirect and through OpenRouter$0.020 in/1M$0.050 out/1M48 tok/sthrough OpenRouter16Kfp86 hours ago direct29 min ago through OpenRouterOpenRouter$0.050 in/1M$0.080 out/1M—131K—30 min agoGroqzero-retentionCHEAPEST$0.050 in/1M$0.080 out/1M124 tok/s131K—29 min agoCoreWeavezero-retention$0.22 in/1M$0.22 out/1M102 tok/s131Kbf1629 min agoCloudflare Workers AI$0.15 in/1M$0.29 out/1M13 tok/s32Kfp829 min ago
Full specs, hardware verdicts and benchmarks on the Llama 3.1 8B Instruct model page →