News / PRICE

DeepSeek V4.1 Flash repriced across 7 hosts: OpenInference input down 70%, InferenceNet input up 97%

Published 29 September 2026
Price move, per 1M tokens
  • InferenceNet · input↑ 97%
    $0.035 → $0.069
  • InferenceNet · output↑ 55%
    $0.29 → $0.45
  • InferenceNet · cache read↑ 1900%
    $0.001 → $0.020
  • OpenInference · input↓ 70%
    $0.100 → $0.030
  • OpenInference · outputunchanged
    $0.50 → $0.50
  • OpenInference · cache readunchanged
    $0.010 → $0.010
  • Morph · inputunchanged
    $0.081 → $0.081
  • Morph · output↑ 64%
    $0.3101 → $0.5100
  • Morph · cache read↑ 74%
    $0.00155 → $0.00270
  • Relace · input↓ 60%
    $0.050 → $0.020
  • Relace · outputunchanged
    $0.60 → $0.60
  • Relace · cache read↑ 100%
    $0.010 → $0.020
  • Ionstream · input↓ 48%
    $0.280 → $0.145
  • Ionstream · outputunchanged
    $1.15 → $1.15
  • Ionstream · cache readunchanged
    $0.005 → $0.005
  • Wafer · input↑ 1%
    $0.099 → $0.100
  • Wafer · output↓ 37%
    $0.70 → $0.44
  • Wafer · cache readunchanged
    $0.045 → $0.045
  • Phala · input↓ 13%
    $0.345 → $0.300
  • Phala · output↓ 13%
    $1.38 → $1.20
  • Phala · cache read↓ 13%
    $0.0069 → $0.0060
InferenceNet input: ↑ 97%OpenInference input: ↓ 70%Morph output: ↑ 64%Relace input: ↓ 60%Relace cache read: ↑ 100%Ionstream input: ↓ 48%Wafer output: ↓ 37%Wafer input: ↑ 1%Phala all rates: ↓ 13%
The model
DeepSeek V4.1 Flashopen the model →
Open weights763.2B parameters1049K context35 hosts
intelligence21st of 168 writing37th of 168 coding11th of 168 agents13th of 55

DeepSeek V4.1 Flash is a downloadable model with a permissive licence, and it places 11th of 168 on Arena Coding via Max as of 25 Sep 2026. At 763.2 billion parameters it is a data-centre job to run yourself, so hosted use is the practical route for most readers.

Source

Read on openrouter.ai · we published this on 29 September 2026.

read the original at openrouter.ai ↗