Models /Llama 3.3 70B Instruct /Where to run
Provider guide

Where to run Llama 3.3 70B Instruct

17 live offers tracked — output prices vary 7.0× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.

CHEAPEST
$0.10 / $0.32
per 1M tokens in / out
FASTEST MEASURED
179 tok/s
measured throughput · $0.59 input / $0.79 output per million tokens
DeepInfra$0.10 in/1M$0.32 out/1M7 tok/s131Kfp86d agoOpenRouterCHEAPEST$0.10 in/1M$0.32 out/1M131K3h agoNovita AI$0.14 in/1M$0.40 out/1M6K9h agoSambaNova$0.45 in/1M$0.90 out/1M103 tok/s131Kbf169h agoSambaNova$0.60 in/1M$1.20 out/1M131K9h agoCloudflare Workers AI$0.29 in/1M$2.25 out/1M43 tok/s24Kfp83h agoDeepInfraturbo tier$0.10 in/1M$0.32 out/1M17 tok/s131Kfp83h agoAkashMLzero-retention$0.13 in/1M$0.40 out/1M21 tok/s131Kfp83h agoNovita AIzero-retention$0.14 in/1M$0.40 out/1M24 tok/s6Kbf163h agoNebius AI Studiozero-retention$0.13 in/1M$0.40 out/1M16 tok/s131Kfp829h agoParasailzero-retention$0.22 in/1M$0.50 out/1M35 tok/s131Kfp823h agoCoreWeavezero-retention$0.71 in/1M$0.71 out/1M71 tok/s128Kfp163h agoGoogle Vertex AIzero-retention$0.72 in/1M$0.72 out/1M53 tok/s128K3h agoGoogle Vertex AIzero-retention$0.72 in/1M$0.72 out/1M51 tok/s128K3h agoCrusoezero-retention$0.25 in/1M$0.75 out/1M60 tok/s131Kbf163h agoGroqzero-retention$0.59 in/1M$0.79 out/1M179 tok/s131K3h agoTogether AIzero-retention$1.04 in/1M$1.04 out/1M42 tok/s131Kfp83h ago
Full specs, hardware verdicts and benchmarks on theLlama 3.3 70B Instruct model page →