Models /Llama 3.1 8B Instruct /Where to run
Provider guide

Where to run Llama 3.1 8B Instruct

6 live listings tracked — output prices vary 7.2× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.

CHEAPEST
$0.050 / $0.080
per 1M tokens in / out
FASTEST MEASURED
124 tok/s
measured throughput · $0.050 input / $0.080 output per million tokens
DeepInfrazero-retention$0.020 in/1M$0.040 out/1M28 tok/s131Kfp82 hours agoNovita AIzero-retention through OpenRouterDirect and through OpenRouter$0.020 in/1M$0.050 out/1M48 tok/sthrough OpenRouter16Kfp88 hours ago direct2 hours ago through OpenRouterOpenRouter$0.050 in/1M$0.080 out/1M—131K—2 hours agoGroqzero-retentionCHEAPEST$0.050 in/1M$0.080 out/1M124 tok/s131K—2 hours agoCoreWeavezero-retention$0.22 in/1M$0.22 out/1M102 tok/s131Kbf162 hours agoCloudflare Workers AI$0.15 in/1M$0.29 out/1M13 tok/s32Kfp82 hours ago
Full specs, hardware verdicts and benchmarks on the Llama 3.1 8B Instruct model page →