DeepSeek V4 Pro
DeepSeek · released Apr 22, 2026 · deepseek-ai/DeepSeek-V4-Pro
- Type
- Open weightsMIT License
- Params
- 1.6T
- Context
- 1M
37B active per word · about 786K words of context
Our take
Written Sep 2, 2026DeepSeek V4 Pro is a large downloadable text model with a one-million-token request limit and strong mathematics scores. Its mixture-of-experts design keeps only 37 billion parameters active per token, making it more efficient to run than its total size suggests.
Choose this for mathematics-heavy workloads or long-document analysis at one million tokens. It suits budget API inference with wide provider choice, or self-hosting under a permissive licence. Skip it if you need image, video or audio input, agentic coding, or consistently fast throughput across providers.
The case for it
- Exceptional mathematics performance on refreshed competition problems: LiveBench Mathematics 90.68%.
- One-million-token request limit — 28.3 times its active parameter count of 37 billion.
- Strong web-development coding score: 122.3 points above its overall Arena Text score.
- Wide provider choice with competitive low-end pricing: 30 offers with a 1.6× spread between cheapest and official pricing.
The case against it
- Weak on agentic tasks: LiveBench Agentic Coding 42.63%, 18.9 points below its next lowest subscore.
- Throughput varies 3.6× by provider, from 16 to 57 tokens per second.
- Text-only: no image, video or audio input.
How good is it?
An open-weights text model for everyday questions and drafting prose.
- getting answers to everyday questionsArena Text (overall) · 36th of 168
- drafts, rewrites and editingArena Creative Writing · 32nd of 168
EverydayGeneral questions and everyday reasoning
Arena Text (overall)36th of 168 · 1464
Also on this board: 1455 (Jul 30, 2026). Read the pair, not the higher one.
CodingWriting and fixing code on its own
Arena Coding43rd of 168 · 1506
Also on this board: 1490 (Jul 30, 2026). Read the pair, not the higher one.
AgenticPlanning, calling tools, staying on task
Arena Agent31st of 55 · −0.011
Also on this board: 0.033 (Sep 25, 2026). Read the pair, not the higher one.
WritingDrafting and rewriting prose
Arena Creative Writing32nd of 168 · 1446
Also on this board: 1442 (Jul 30, 2026). Read the pair, not the higher one.
Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.
Boards this model appears on that none of the ratings above are built on.
Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Apple M1 Pro (16-core GPU) · 32 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Radeon RX 7900 XT · 20 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 4 hours and 46 hours ago — each listing carries its own date.
Some hosts sell this model at two prices: on their own price list (“direct”) and on their OpenRouter listing (“through OpenRouter”). Where the two differ, the row shows both, each with the date we last read it.
- per 1M tokens
- $0.78 in / $1.57 out
- Context served
- 1M
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| OpenRouterOpenRouter's own listing | $0.78 / $1.57checked 22 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| StreamLakefp8Through OpenRouter | $0.78 / $1.57checked 22 hours ago | 1M384K max reply | 41 tok/s | No | Yesunknown period | Unknown |
| Alibaba CloudThrough OpenRouter | $0.58 / $1.74checked 4 hours ago | 1M393K max reply | 63 tok/s | No | Yesunknown period | Unknown |
| RekaThrough OpenRouter | $0.58 / $1.74checked 4 hours ago | 1M393K max reply | 28 tok/s | No | No | Confirmed |
| IonstreamThrough OpenRouter | $1.24 / $1.85checked 46 hours ago | 1M944K max reply | 82 tok/s | No | No | Confirmed |
| GMICloudfp8Through OpenRouter | $0.96 / $1.91checked 4 hours ago | 1M944K max reply | 23 tok/s | No | Yesunknown period | Unknown |
| DeepSeekThrough OpenRouter | $0.66 / $1.98checked 4 hours ago | 1M393K max reply | 9 tok/s | Yes | Yesunknown period | Unknown |
| StreamLakeThrough OpenRouter | $0.66 / $1.98checked 4 hours ago | 1M384K max reply | 45 tok/s | No | Yesunknown period | Unknown |
| DigitalOcean GradientThrough OpenRouter | $1.04 / $2.09checked 4 hours ago | 1M384K max reply | 49 tok/s | No | No | Confirmed |
| Cloudflare Workers AIThrough OpenRouter | $1.15 / $2.55checked 4 hours ago | 1M944K max reply | 39 tok/s | No | Yesunknown period | Unknown |
| DeepInfrafp8Direct and through OpenRouter | $1.30 / $2.60checked 4 hours ago | 1M16K max reply through OpenRouter | 33 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| Alibaba Cloudfp8Through OpenRouter | $1.42 / $2.83checked 10 hours ago | 1M393K max reply | 41 tok/s | No | Yesunknown period | Unknown |
| PhalaThrough OpenRouter | $0.96 / $2.88checked 4 hours ago | 1M393K max reply | 49 tok/s | No | No | Confirmed |
| Novita AIfp8Direct and through OpenRouter | $1.60 / $3.20directchecked 4 hours ago$0.99 / $2.97through OpenRouterchecked 4 hours ago | 1M393K max reply through OpenRouter | 49 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| SiliconFlowfp8Through OpenRouter | $1.50 / $3.13checked 4 hours ago | 1M393K max reply | 37 tok/s | No | No | Confirmed |
| NextBitfp8Through OpenRouter | $1.06 / $3.17checked 4 hours ago | 1M944K max reply | 50 tok/s | No | No | Confirmed |
| Venice AIThrough OpenRouter | $1.65 / $3.30checked 4 hours ago | 1M33K max reply | 51 tok/s | No | No | Confirmed |
| AtlasCloudfp4Through OpenRouter | $1.68 / $3.38checked 4 hours ago | 1M393K max reply | 45 tok/s | No | Yesunknown period | Unknown |
| Relacefp4Through OpenRouter | $0.17 / $3.50checked 4 hours ago | 1M393K max reply | 75 tok/s | No | No | Confirmed |
| Microsoft Azure AIusThrough OpenRouter | $1.91 / $3.83checked 4 hours ago | 1M384K max reply | 58 tok/s | No | No | Confirmed |
| AtlasCloudfp8Through OpenRouter | $1.32 / $3.96checked 4 hours ago | 1M393K max reply | 34 tok/s | No | Yesunknown period | Unknown |
| Baidufp8Through OpenRouter | $1.32 / $3.96checked 40 hours ago | 1M393K max reply | 45 tok/s | No | Yesunknown period | Unknown |
| Parasailfp8Through OpenRouter | $1.32 / $3.96checked 4 hours ago | 1M944K max reply | 65 tok/s | No | No | Confirmed |
| CoreWeavefp8Through OpenRouter | $1.31 / $3.96checked 4 hours ago | 1M944K max reply | 60 tok/s | No | No | Confirmed |
| Together AIThrough OpenRouter | $1.32 / $3.96checked 16 hours ago | 1M944K max reply | 110 tok/s | No | No | Confirmed |
| WaferThrough OpenRouter | $0.55 / $4.20checked 4 hours ago | 1M944K max reply | 48 tok/s | No | No | Confirmed |
Across the 26 listings we hold: 24 say they do not train on prompts (2 of them only through OpenRouter), 1 says it does and 1 does not say. 15 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| StreamLakefp8Through OpenRouter | ✓ | ✓ | ✓ |
| Alibaba CloudThrough OpenRouter | ✓ | ✓ | ✓ |
| RekaThrough OpenRouter | ✓ | ✓ | ✓ |
| IonstreamThrough OpenRouter | ✓ | ✓ | ✓ |
| GMICloudfp8Through OpenRouter | ✓ | ✓ | ✗ |
| DeepSeekThrough OpenRouter | ✓ | ✓ | ✗ |
| StreamLakeThrough OpenRouter | ✓ | ✓ | ✗ |
| DigitalOcean GradientThrough OpenRouter | ✓ | ✓ | ✗ |
| Cloudflare Workers AIThrough OpenRouter | ✓ | ✓ | ✓ |
| DeepInfrafp8Direct and through OpenRouter | ✓ | ✓ | ✓ |
| Alibaba Cloudfp8Through OpenRouter | ✓ | ✓ | ✓ |
| PhalaThrough OpenRouter | ✓ | ✓ | ✓ |
| Novita AIfp8Direct and through OpenRouter | ✓ | ✓ | ✗ |
| SiliconFlowfp8Through OpenRouter | ✓ | ✓ | ✗ |
| NextBitfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Venice AIThrough OpenRouter | ✓ | ✓ | ✓ |
| AtlasCloudfp4Through OpenRouter | ✓ | ✓ | ✓ |
| Relacefp4Through OpenRouter | ✓ | ✓ | ✗ |
| Microsoft Azure AIusThrough OpenRouter | ✓ | ✓ | ✗ |
| AtlasCloudfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Baidufp8Through OpenRouter | ✓ | ✓ | ✓ |
| Parasailfp8Through OpenRouter | ✓ | ✓ | ✓ |
| CoreWeavefp8Through OpenRouter | ✓ | ✓ | ✓ |
| Together AIThrough OpenRouter | ✓ | ✓ | ✓ |
| WaferThrough OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 26 of 26 listings say yes. JSON output: 26 of 26 listings say yes. Strict schema: 18 of 26 listings say yes, 8 say no.
Models people weigh against DeepSeek V4 Pro
When we formed this view
Recent changes
What moved
DeepSeek V4 Pro moved on 3 hosts: Reka: input −55% ($1.30 → $0.58 per 1M tokens), output −33% ($2.60 → $1.74 per 1M tokens), cache read −55% ($0.130 → $0.058 per 1M tokens); Relace: input +36% ($0.123 → $0.167 per 1M tokens), cache read +36% ($0.123 → $0.167 per 1M tokens); Wafer: input +15% ($0.48 → $0.55 per 1M tokens), cache read +15% ($0.384 → $0.440 per 1M tokens)What moved
DeepSeek V4 Pro moved on 3 hosts: Ionstream: input +104% ($0.608 → $1.238 per 1M tokens), output −6% ($1.96 → $1.85 per 1M tokens); Relace: input −18% ($0.150 → $0.123 per 1M tokens), cache read −18% ($0.150 → $0.123 per 1M tokens); Wafer: input −1% ($0.480 → $0.475 per 1M tokens), output −17% ($4.20 → $3.49 per 1M tokens), cache read +24% ($0.38 → $0.47 per 1M tokens)What moved
DeepSeek V4 Pro moved on 5 hosts: Ionstream: input +169% ($0.226 → $0.608 per 1M tokens); Relace: input −57% ($0.35 → $0.15 per 1M tokens), cache read +50% ($0.10 → $0.15 per 1M tokens); Io Net: input −29% ($0.99 → $0.70 per 1M tokens), output +11% ($3.15 → $3.50 per 1M tokens), cache read −17% ($0.108 → $0.090 per 1M tokens); GMICloud: input −24% ($1.74 → $1.32 per 1M tokens), output +14% ($3.48 → $3.96 per 1M tokens), cache read −70% ($0.145 → $0.044 per 1M tokens); Wafer: input −0.2% ($0.395 → $0.394 per 1M tokens), output −17% ($4.20 → $3.49 per 1M tokens), cache read −0.2% ($0.316 → $0.315 per 1M tokens)What moved
DeepSeek V4 Pro moved on 4 hosts: Relace: input +72% ($0.204 → $0.350 per 1M tokens), output −7% ($3.78 → $3.50 per 1M tokens), cache read −51% ($0.203 → $0.100 per 1M tokens); Wafer: input +35% ($0.293 → $0.397 per 1M tokens), output +20% ($3.50 → $4.20 per 1M tokens), cache read +85% ($0.172 → $0.318 per 1M tokens); Io Net: input −11% ($1.24 → $1.10 per 1M tokens), cache read −14% ($0.14 → $0.12 per 1M tokens); Ionstream: input −11% ($0.253 → $0.226 per 1M tokens)What moved
input +20% ($0.245 → $0.293 per 1M tokens), cache read −30% ($0.245 → $0.172 per 1M tokens)What moved
DeepSeek V4 Pro moved on 2 hosts: Wafer: input −36% ($0.39 → $0.25 per 1M tokens), output +21% ($2.90 → $3.50 per 1M tokens), cache read −0.4% ($0.250 → $0.249 per 1M tokens); Ionstream: input −35% ($0.390 → $0.253 per 1M tokens), output −32% ($2.88 → $1.96 per 1M tokens)Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 1 of 26 listings does not say whether it trains on prompts, and 2 answer only through OpenRouter, not for their own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsMIT License, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
MIT License
Fully permissive: do anything with attribution. No patent grant, unlike Apache-2.0.
Identifiers
- Hugging Face
- deepseek-ai/DeepSeek-V4-Pro
- Architecture
- Mixture of experts
- Takes in, gives back
- Text in, text out
- Catalogue slug
- deepseek-deepseek-v4-pro