DeepSeek V4.1 Flash repriced across 7 hosts: OpenInference input down 70%, InferenceNet input up 97%
Published 29 September 2026
Price move, per 1M tokens
- InferenceNet · input↑ 97%$0.035 → $0.069
- InferenceNet · output↑ 55%$0.29 → $0.45
- InferenceNet · cache read↑ 1900%$0.001 → $0.020
- OpenInference · input↓ 70%$0.100 → $0.030
- OpenInference · outputunchanged$0.50 → $0.50
- OpenInference · cache readunchanged$0.010 → $0.010
- Morph · inputunchanged$0.081 → $0.081
- Morph · output↑ 64%$0.3101 → $0.5100
- Morph · cache read↑ 74%$0.00155 → $0.00270
- Relace · input↓ 60%$0.050 → $0.020
- Relace · outputunchanged$0.60 → $0.60
- Relace · cache read↑ 100%$0.010 → $0.020
- Ionstream · input↓ 48%$0.280 → $0.145
- Ionstream · outputunchanged$1.15 → $1.15
- Ionstream · cache readunchanged$0.005 → $0.005
- Wafer · input↑ 1%$0.099 → $0.100
- Wafer · output↓ 37%$0.70 → $0.44
- Wafer · cache readunchanged$0.045 → $0.045
- Phala · input↓ 13%$0.345 → $0.300
- Phala · output↓ 13%$1.38 → $1.20
- Phala · cache read↓ 13%$0.0069 → $0.0060
| host | rate | was | now | change |
|---|---|---|---|---|
| InferenceNet | input | $0.035 | $0.069 | ↑ 97% |
| InferenceNet | output | $0.29 | $0.45 | ↑ 55% |
| InferenceNet | cache read | $0.001 | $0.020 | ↑ 1900% |
| OpenInference | input | $0.100 | $0.030 | ↓ 70% |
| OpenInference | output | $0.50 | $0.50 | — |
| OpenInference | cache read | $0.010 | $0.010 | — |
| Morph | input | $0.081 | $0.081 | — |
| Morph | output | $0.3101 | $0.5100 | ↑ 64% |
| Morph | cache read | $0.00155 | $0.00270 | ↑ 74% |
| Relace | input | $0.050 | $0.020 | ↓ 60% |
| Relace | output | $0.60 | $0.60 | — |
| Relace | cache read | $0.010 | $0.020 | ↑ 100% |
| Ionstream | input | $0.280 | $0.145 | ↓ 48% |
| Ionstream | output | $1.15 | $1.15 | — |
| Ionstream | cache read | $0.005 | $0.005 | — |
| Wafer | input | $0.099 | $0.100 | ↑ 1% |
| Wafer | output | $0.70 | $0.44 | ↓ 37% |
| Wafer | cache read | $0.045 | $0.045 | — |
| Phala | input | $0.345 | $0.300 | ↓ 13% |
| Phala | output | $1.38 | $1.20 | ↓ 13% |
| Phala | cache read | $0.0069 | $0.0060 | ↓ 13% |
InferenceNet input: ↑ 97%OpenInference input: ↓ 70%Morph output: ↑ 64%Relace input: ↓ 60%Relace cache read: ↑ 100%Ionstream input: ↓ 48%Wafer output: ↓ 37%Wafer input: ↑ 1%Phala all rates: ↓ 13%
The model
DeepSeek V4.1 Flashopen the model →
Open weights763.2B parameters1049K context35 hosts
intelligence21st of 168 writing37th of 168 coding11th of 168 agents13th of 55
DeepSeek V4.1 Flash is a downloadable model with a permissive licence, and it places 11th of 168 on Arena Coding via Max as of 25 Sep 2026. At 763.2 billion parameters it is a data-centre job to run yourself, so hosted use is the practical route for most readers.
Source
Read on openrouter.ai · we published this on 29 September 2026.
read the original at openrouter.ai ↗