DeepSeek V4 Flash
DeepSeek · released Apr 22, 2026 · deepseek-ai/DeepSeek-V4-Flash
- Type
- Open weightsMIT License
- Params
- 291B
- Context
- 1M
37B active per word · about 786K words of context
Our take
Written Sep 30, 2026DeepSeek V4 Flash is a downloadable text model with a permissive licence, built for long-document work where the bill matters more than peak quality. Its measured quality is weak on the LiveBench boards and mid-pack on the Arena boards, so treat it as a cheap workhorse rather than a quality leader.
Use it for long-document work where the bill matters more than peak quality, and where the request capacity means long documents need not be split up first. You can run it yourself on a single modern graphics card, and the licence allows commercial use, changes and redistribution. Skip it when a task needs measured coding or reasoning evidence at a competitive level, or when you need serving speed figures to choose a host.
The case for it
- 37 billion of its 290.9 billion parameters work per token, so memory in use is closer to a small model than to a mid-size one, making it realistic to run on one machine.
- The MIT License allows commercial use, changes and redistribution, so the terms stay out of the way of a commercial product.
- The request capacity leaves room for a long report or a stack of documents beside the question, though reliable recall across all of it is unverified in our data.
The case against it
- 55th of 58 on LiveBench as of 25 Jun 2026, and 57th of 58 on LiveBench Agentic Coding as of 25 Jun 2026, so it sits near the bottom of the field on those tasks.
- 65th of 168 on Arena Text (overall) as of 25 Sep 2026, and 64th of 168 on Arena Coding as of 25 Sep 2026, so its Arena standings are mid-pack rather than leading.
- Rates differ between the listed hosts, and no serving speed is supplied for any of them, so price alone cannot pick the host.
How good is it?
EverydayGeneral questions and everyday reasoning
Arena Text (overall)65th of 168 · 1436
Also on this board: 1438 (Jul 30, 2026). Read the pair, not the higher one.
CodingWriting and fixing code on its own
Arena Coding64th of 168 · 1484
Also on this board: 1481 (Jul 30, 2026). Read the pair, not the higher one.
AgenticPlanning, calling tools, staying on task
Arena Agent20th of 55 · 0.018
WritingDrafting and rewriting prose
Arena Creative Writing57th of 168 · 1408
Also on this board: 1406 (Jul 30, 2026). Read the pair, not the higher one.
Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.
Boards this model appears on that none of the ratings above are built on.
Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Apple M1 Pro (16-core GPU) · 32 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Comfortable fit
Apple M3 Ultra (80-core GPU) · 512 GB
Room to spare. 193.7 GB spare means a 10% error in the size would not change the answer.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 4 hours and 7 days ago — each listing carries its own date.
- per 1M tokens
- $0.042 in / $0.084 out
- Context served
- 1M
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| OpenRouterOpenRouter's own listing | $0.042 / $0.084checked 4 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| StreamLakefp8Through OpenRouter | $0.042 / $0.084checked 4 hours ago | 1M384K max reply | 60 tok/s | No | Yesunknown period | Unknown |
| DeepInfrafp8Direct and through OpenRouter | $0.060 / $0.18checked 4 hours ago | 1M384K max reply through OpenRouter | 15 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| GMICloudfp8Through OpenRouter | $0.091 / $0.18checked 4 hours ago | 1M944K max reply | 52 tok/s | No | Yesunknown period | Unknown |
| Venice AIThrough OpenRouter | $0.097 / $0.19checked 4 hours ago | 1M33K max reply | 19 tok/s | No | No | Confirmed |
| MakoraThrough OpenRouter | $0.090 / $0.20checked 4 hours ago | 1M384K max reply | 35 tok/s | No | No | Confirmed |
| DigitalOcean GradientThrough OpenRouter | $0.098 / $0.20checked 4 hours ago | 1M384K max reply | 14 tok/s | No | No | Confirmed |
| Basetenfp8Through OpenRouter | $0.13 / $0.26checked 4 hours ago | 1M384K max reply | 71 tok/s | No | No | Confirmed |
| Alibaba Cloudfp8Through OpenRouter | $0.13 / $0.27checked 4 hours ago | 1M393K max reply | 86 tok/s | No | Yesunknown period | Unknown |
| Novita AIfp8Direct and through OpenRouter | $0.14 / $0.28checked 4 hours ago | 1M393K max reply through OpenRouter | 63 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| SiliconFlowfp8Through OpenRouter | $0.13 / $0.28checked 4 hours ago | 1M393K max reply | 24 tok/s | No | No | Confirmed |
| AtlasCloudfp4Through OpenRouter | $0.14 / $0.28checked 4 hours ago | 1M393K max reply | 62 tok/s | No | Yesunknown period | Unknown |
| Parasailfp8Through OpenRouter | $0.14 / $0.28checked 4 hours ago | 1M944K max reply | 71 tok/s | No | No | Confirmed |
| CoreWeavefp8Through OpenRouter | $0.13 / $0.28checked 4 hours ago | 262K236K max reply | 58 tok/s | No | No | Confirmed |
| CohereThrough OpenRouter | $0.14 / $0.28checked 4 hours ago | 1M384K max reply | 94 tok/s | No | Yes30 days | Confirmed |
| Together AIThrough OpenRouter | $0.14 / $0.28checked 4 hours ago | 1M944K max reply | 30 tok/s | No | No | Confirmed |
| Sail Researchfp4Through OpenRouter | $0.019 / $0.30checked 4 hours ago | 1M944K max reply | 31 tok/s | No | No | Confirmed |
| Morphbf16Through OpenRouter | $0.14 / $0.40checked 4 hours ago | 1M944K max reply | 38 tok/s | No | No | Confirmed |
| Sail Researchusfp4Through OpenRouter | $0.019 / $0.42checked 4 hours ago | 1M944K max reply | 33 tok/s | No | No | Unknown |
| Mancer 2fp8Through OpenRouter | $0.19 / $0.50checked 4 hours ago | 1M944K max reply | 19 tok/s | No | No | Confirmed |
| RekaThrough OpenRouter | $0.088 / $0.53checked 4 days ago | 262K131K max reply | 105 tok/s | No | No | Confirmed |
| Alibaba CloudThrough OpenRouter | $0.18 / $0.53checked 4 hours ago | 1M393K max reply | 55 tok/s | No | Yesunknown period | Unknown |
| Microsoft Azure AIusThrough OpenRouter | $0.21 / $0.56checked 4 hours ago | 1M384K max reply | 60 tok/s | No | No | Confirmed |
| Inceptronfp4Through OpenRouter | $0.050 / $0.65checked 4 hours ago | 1M944K max reply | 36 tok/s | No | No | Confirmed |
| OpenInferencefp8Through OpenRouter | $0.14 / $0.70checked 7 days ago | 1M944K max reply | 27 tok/s | No | No | Confirmed |
| Waferfast tierThrough OpenRouter | $0.12 / $0.70checked 4 hours ago | 1M944K max reply | 77 tok/s | No | No | Confirmed |
| PhalaThrough OpenRouter | $0.31 / $0.92checked 4 hours ago | 1M393K max reply | 43 tok/s | No | No | Confirmed |
| NextBitfp8Through OpenRouter | $0.35 / $1.06checked 4 hours ago | 1M944K max reply | 54 tok/s | No | No | Confirmed |
| Relacefp4Through OpenRouter | $0.004 / $1.28checked 4 hours ago | 1M944K max reply | 65 tok/s | No | No | Confirmed |
| Cloudflare Workers AIThrough OpenRouter | $0.44 / $1.32checked 4 hours ago | 1M944K max reply | 54 tok/s | No | Yesunknown period | Unknown |
| Baidufp8Through OpenRouter | $0.44 / $1.32checked 28 hours ago | 1M131K max reply | 96 tok/s | No | Yesunknown period | Unknown |
Across the 31 listings we hold: 30 say they do not train on prompts (2 of them only through OpenRouter), 0 say they do and 1 does not say. 22 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| StreamLakefp8Through OpenRouter | ✓ | ✓ | ✓ |
| DeepInfrafp8Direct and through OpenRouter | ✓ | ✓ | ✓ |
| GMICloudfp8Through OpenRouter | ✓ | ✓ | ✗ |
| Venice AIThrough OpenRouter | ✓ | ✓ | ✓ |
| MakoraThrough OpenRouter | ✓ | ✓ | ✓ |
| DigitalOcean GradientThrough OpenRouter | ✓ | ✓ | ✓ |
| Basetenfp8Through OpenRouter | ✓ | ✗ | ✗ |
| Alibaba Cloudfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Novita AIfp8Direct and through OpenRouter | ✓ | ✓ | ✗ |
| SiliconFlowfp8Through OpenRouter | ✓ | ✓ | ✗ |
| AtlasCloudfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Parasailfp8Through OpenRouter | ✓ | ✓ | ✓ |
| CoreWeavefp8Through OpenRouter | ✓ | ✗ | ✗ |
| CohereThrough OpenRouter | ✓ | ✓ | ✓ |
| Together AIThrough OpenRouter | ✓ | ✓ | ✓ |
| Sail Researchfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Morphbf16Through OpenRouter | ✓ | ✓ | ✓ |
| Sail Researchus · fp4Through OpenRouter | ✓ | ✓ | ✓ |
| Mancer 2fp8Through OpenRouter | ✓ | ✓ | ✓ |
| RekaThrough OpenRouter | ✓ | ✓ | ✓ |
| Alibaba CloudThrough OpenRouter | ✓ | ✓ | ✓ |
| Microsoft Azure AIusThrough OpenRouter | ✓ | ✓ | ✗ |
| Inceptronfp4Through OpenRouter | ✓ | ✓ | ✓ |
| OpenInferencefp8Through OpenRouter | ✓ | ✓ | ✓ |
| WaferfastThrough OpenRouter | ✓ | ✓ | ✓ |
| PhalaThrough OpenRouter | ✓ | ✓ | ✓ |
| NextBitfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Relacefp4Through OpenRouter | ✓ | ✓ | ✗ |
| Cloudflare Workers AIThrough OpenRouter | ✓ | ✓ | ✓ |
| Baidufp8Through OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 31 of 31 listings say yes. JSON output: 29 of 31 listings say yes, 2 say no. Strict schema: 24 of 31 listings say yes, 7 say no.
Models people weigh against DeepSeek V4 Flash
When we formed this view
Recent changes
What moved
DeepSeek V4 Flash moved on 3 hosts: Wafer: input +71% ($0.070 → $0.120 per 1M tokens), output +100% ($0.35 → $0.70 per 1M tokens), cache read −6% ($0.053 → $0.050 per 1M tokens); StreamLake: input −68% ($0.44 → $0.14 per 1M tokens), output −79% ($1.32 → $0.28 per 1M tokens), cache read +100% ($0.014 → $0.028 per 1M tokens); Relace: input −50% ($0.0090 → $0.0045 per 1M tokens), cache read −50% ($0.0090 → $0.0045 per 1M tokens)What moved
DeepSeek V4 Flash moved on 2 hosts: Relace: input −50% ($0.018 → $0.009 per 1M tokens), output +300% ($0.32 → $1.28 per 1M tokens), cache read −50% ($0.018 → $0.009 per 1M tokens); Inceptron: input −11% ($0.056 → $0.050 per 1M tokens)What moved
DeepSeek V4 Flash moved on 5 hosts: Inceptron: input −14% ($0.065 → $0.056 per 1M tokens), output +71% ($0.38 → $0.65 per 1M tokens); AtlasCloud: input −48% ($0.1584 → $0.0826 per 1M tokens), output −65% ($0.48 → $0.17 per 1M tokens), cache read +64% ($0.01008 → $0.01652 per 1M tokens); Venice: input +27% ($0.138 → $0.175 per 1M tokens), output +27% ($0.275 → $0.350 per 1M tokens), cache read +25% ($0.028 → $0.035 per 1M tokens); Wafer: input +19% ($0.059 → $0.070 per 1M tokens); Relace: input −14% ($0.021 → $0.018 per 1M tokens), cache read +12% ($0.016 → $0.018 per 1M tokens)What moved
DeepSeek V4 Flash moved on 5 hosts: NextBit: input +135% ($0.150 → $0.352 per 1M tokens), output +252% ($0.300 → $1.056 per 1M tokens), cache read −66% ($0.035 → $0.012 per 1M tokens); Wafer: input +106% ($0.0413 → $0.0850 per 1M tokens), cache read +43% ($0.037 → $0.053 per 1M tokens); AtlasCloud: input +13% ($0.140 → $0.158 per 1M tokens), output +70% ($0.280 → $0.475 per 1M tokens), cache read −64% ($0.028 → $0.010 per 1M tokens); Inceptron: input +67% ($0.039 → $0.065 per 1M tokens), output +40% ($0.271 → $0.380 per 1M tokens), cache read −10% ($0.030 → $0.027 per 1M tokens); Sail Research (US region): input −16% ($0.0225 → $0.0190 per 1M tokens); Sail Research: input −12% ($0.0215 → $0.0190 per 1M tokens)What moved
input −30% ($0.0590 → $0.0413 per 1M tokens), cache read −36% ($0.058 → $0.037 per 1M tokens)What moved
DeepSeek V4 Flash moved on 3 hosts: Inceptron: input −53% ($0.083 → $0.039 per 1M tokens), output −44% ($0.48 → $0.27 per 1M tokens), cache read −40% ($0.050 → $0.030 per 1M tokens); Sail Research: input −28% ($0.0300 → $0.0215 per 1M tokens), output −45% ($0.55 → $0.30 per 1M tokens), cache read −13% ($0.016 → $0.014 per 1M tokens); Relace: input −30% ($0.030 → $0.021 per 1M tokens); Sail Research (US region): input −25% ($0.0300 → $0.0225 per 1M tokens), output −24% ($0.55 → $0.42 per 1M tokens), cache read −25% ($0.016 → $0.012 per 1M tokens)What moved
DeepSeek V4 Flash moved on 2 hosts: Sail Research: input −21% ($0.038 → $0.030 per 1M tokens), cache read −30% ($0.023 → $0.016 per 1M tokens); Sail Research (US region): input −21% ($0.038 → $0.030 per 1M tokens), cache read −30% ($0.023 → $0.016 per 1M tokens); Inceptron: output +17% ($0.41 → $0.48 per 1M tokens), cache read −17% ($0.060 → $0.050 per 1M tokens)Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 1 of 31 listings does not say whether it trains on prompts, and 2 answer only through OpenRouter, not for their own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsMIT License, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
MIT License
Fully permissive: do anything with attribution. No patent grant, unlike Apache-2.0.
Identifiers
- Hugging Face
- deepseek-ai/DeepSeek-V4-Flash
- Architecture
- Mixture of experts
- Takes in, gives back
- Text in, text out
- Catalogue slug
- deepseek-deepseek-v4-flash