Models /DeepSeek V4.1 Flash /Where to run
Provider guide

Where to run DeepSeek V4.1 Flash

33 live listings tracked — output prices vary 3.8× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.

CHEAPEST
$0.12 / $0.40
per 1M tokens in / out
FASTEST MEASURED
221 tok/s
measured throughput · $0.30 input / $1.20 output per million tokens
Sail Researchzero-retention$0.080 in/1M$0.40 out/1M42 tok/s1Mfp41 hour agoDekaLLMzero-retentionCHEAPEST$0.12 in/1M$0.40 out/1M43 tok/s1M—1 hour agoIo Netzero-retention$0.12 in/1M$0.41 out/1M94 tok/s1Mfp81 hour agoDeepInfrazero-retention through OpenRouterDirect and through OpenRouter$0.20 in/1M$0.14 in/1M$0.60 out/1Mdirect$0.42 out/1Mthrough OpenRouter74 tok/sthrough OpenRouter1Mfp87 hours ago direct1 hour ago through OpenRouterMorphzero-retention$0.053 in/1M$0.43 out/1M97 tok/s1Mfp81 hour agoStreamLake$0.14 in/1M$0.56 out/1M88 tok/s1Mfp81 hour agoRekazero-retention$0.14 in/1M$0.56 out/1M88 tok/s1M—1 hour agoAtlasCloud$0.14 in/1M$0.56 out/1M49 tok/s1Mfp81 hour agoDeepSeek$0.15 in/1M$0.60 out/1M82 tok/s1M—13 hours agoRelacezero-retention$0.023 in/1M$0.60 out/1M48 tok/s1M—1 hour agoWaferzero-retention$0.050 in/1M$0.60 out/1M16 tok/s1M—1 hour agoInferenceNetzero-retention$0.040 in/1M$0.60 out/1M17 tok/s1M—1 hour agoOpenInferencezero-retention$0.026 in/1M$0.63 out/1M13 tok/s1Mfp41 hour agoCoreWeavezero-retention$0.20 in/1M$0.65 out/1M147 tok/s1Mfp81 hour agoGMICloud$0.18 in/1M$0.72 out/1M109 tok/s1Mfp81 hour agoDigitalOcean Gradientzero-retention$0.18 in/1M$0.72 out/1M67 tok/s1M—1 hour agoNextBitzero-retention$0.21 in/1M$0.84 out/1M107 tok/s1Mfp81 hour agoPhalazero-retention$0.21 in/1M$0.84 out/1M82 tok/s1M—1 hour agoMakorazero-retention$0.20 in/1M$0.99 out/1M189 tok/s1Mfp81 hour agoIonstreamzero-retention$0.28 in/1M$1.15 out/1M157 tok/s1M—43 hours agoBaidu$0.30 in/1M$1.20 out/1M159 tok/s1Mfp81 hour agoParasailzero-retention$0.30 in/1M$1.20 out/1M136 tok/s1Mfp81 hour agoOpenRouter$0.30 in/1M$1.20 out/1M—1M—43 hours agoBasetenzero-retention$0.30 in/1M$1.20 out/1M97 tok/s1Mfp81 hour agoFireworks AIzero-retention$0.30 in/1M$1.20 out/1M72 tok/s1M—1 hour agoAlibaba Cloud$0.30 in/1M$1.20 out/1M81 tok/s1M—1 hour agoModalzero-retention$0.30 in/1M$1.20 out/1M93 tok/s1M—1 hour agoNovita AIzero-retention through OpenRouterDirect and through OpenRouter$0.30 in/1M$1.20 out/1M102 tok/sthrough OpenRouter1Mfp87 hours ago direct1 hour ago through OpenRouterSiliconFlowzero-retention$0.30 in/1M$1.20 out/1M78 tok/s1Mfp81 hour agoTogether AIzero-retention$0.30 in/1M$1.20 out/1M221 tok/s1M—1 hour agoVenice AIzero-retention$0.38 in/1M$1.50 out/1M147 tok/s1Mfp81 hour agoFireworks AIzero-retention$0.45 in/1M$1.80 out/1M212 tok/s1M—1 hour agoBasetenfast tier$0.60 in/1M$2.40 out/1M32 tok/s1Mfp321 hour ago
Full specs, hardware verdicts and benchmarks on the DeepSeek V4.1 Flash model page →