Models / DeepSeek/ DeepSeek V4 Pro

DeepSeek V4 Pro

DeepSeek · released Apr 22, 2026 · deepseek-ai/DeepSeek-V4-Pro

Input: text. Output: text.InputOutput
Type
Open weightsMIT License
Params
1.6T
Context
1M

37B active per word · about 786K words of context

Our take

Written Aug 3, 2026

DeepSeek V4 Pro is a large downloadable text model with a one-million-token request limit and a permissive MIT licence. It excels at coding and hard-prompt tasks on the independent leaderboards we track, though its creative-writing score lags behind its other skills.

Who should pick it

Choose this for coding-heavy workloads where its highest arena score matters, or for long-document analysis at one million tokens. It suits cost-sensitive inference through the cheapest hosts, and teams who need a permissive licence for commercial use or modification. Skip it if you need image, video or audio input, if creative writing is your main use, or if you want predictable speed across providers.

The case for it

  • Exceptional coding performance relative to its other capabilities: its coding score is 44 points above its overall text score on the arena leaderboard.
  • One-million-token request limit, rare among downloadable models, for long-document analysis.
  • Strong hard-prompt handling, scoring 23 points above its overall text score.
  • Permissive MIT licence allows commercial use, modification and redistribution.

The case against it

  • Creative writing is its weakest arena skill, scoring 56 points below its coding score and 12 points below its overall text score.
  • Throughput varies more than sixfold across providers, so speed depends heavily on which host you pick.
  • Text-only: no image, video or audio input or output.
00

How good is it?

IntelligencePuzzles, maths, exam questions

3.5 of 5

Arena Text (overall)28th of 143 · 1457.9

Arena Hard Prompts 28th of 143Arena Maths 36th of 139LiveBench Mathematics 13th of 35LiveBench Data Analysis 14th of 35LiveBench Reasoning 20th of 35

Also on this board: 1455.1 via Thinking (Jul 30, 2026). Read the pair, not the higher one.

CodingWriting and fixing code on its own

3.5 of 5

Arena Coding32nd of 143 · 1502.2

Arena Code (WebDev) 31st of 74LiveBench Coding 30th of 35

Also on this board: 1489.9 via Thinking (Jul 30, 2026). Read the pair, not the higher one.

AgenticPlanning, calling tools, staying on task

2.5 of 5

Arena Agent (IPS)21st of 36 · −0.008

LiveBench Agentic Coding 27th of 35

WritingWe do not rate this

Scored, not ratedThe placings are on the right.

Two boards come close and neither tests writing: Arena Creative Writing asks people which of two replies they prefer, and LiveBench Language tests whether a model understood a passage. So we show where DeepSeek V4 Pro placed and give it no mark out of five.

Arena Creative Writing 22nd of 143 · 1446.1LiveBench Language 16th of 35 · 78.1
Also scored, on boards we give no mark for
Arena Instruction Following 25th of 143LiveBench 24th of 35LiveBench Instruction Following 27th of 35

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.

Every published score for this model16 scoresEvery figure we hold, from 16 boards, with who ran it and a link to the source — including the boards no rating above is built on.
LiveBenchreasoning
71.6independentsource ↗
70independentsource ↗
74.5independentsource ↗
78.1independentsource ↗
82.7independentsource ↗
−0.008independentsource ↗
1502.2independentsource ↗
1480.9independentsource ↗
1445.3independentsource ↗
1457.9independentsource ↗
1446independentsource ↗
01

Can you run it yourself?

Fits in memory
weights load entirely on the card
Spills to system RAM
some weights offload; much slower
Too large
will not load even with offload
est
size is calculated; the verdict could change by 10%
A card many people ownToo large

GeForce RTX 4090 · 24 GB

Weights at Q4_K_M1008 / 24 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

One step upToo large

Apple M1 Pro (16-core GPU) · 32 GB

Weights at Q4_K_M1008 / 32 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

One step downToo large

Radeon RX 7900 XT · 20 GB

Weights at Q4_K_M1008 / 20 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

Your hardware
Checking your profile…

Memory use by level

Against a 24 GB card.

Q4_K_M
recommended
1008 GBest
Too large
Q5_K_M
1182.6 GBest
Too large
Q8_0
1766.7 GBest
Too large

All 3 sizes here are calculated, not measured. We hold no measured file for this model, so each size comes from the parameter count and every verdict above inherits that uncertainty.

Check against your own machine → · All 71 devices, with every size →

02

Or rent it from someone else

Cheapest published offer

Cheapest of 23 live listings. Picked at the widest standard context we hold, within one quantisation slice, so the numbers beside it are a price one host actually charges.

per 1M tokens
$0.43 in / $0.87 out
Context served
1M
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
OpenRouter$0.43 / $0.871Mnot measuredUnknownUnknownUnknown
DeepSeek$0.43 / $0.871M27 tok/sYesYesunknown periodUnknown
Baidufp8$0.54 / $1.081M61 tok/sNoYesunknown periodUnknown
StreamLakefp8$0.61 / $1.221M32 tok/sNoYesunknown periodUnknown
GMICloudfp8$0.68 / $1.361M29 tok/sNoYesunknown periodUnknown
Ionstreamfp4$1.13 / $2.261M14 tok/sNoNoConfirmed
Novita AIfp8$1.17 / $2.341M38 tok/sNoNoConfirmed
Waferfp4$1.20 / $2.401Mnot measuredNoNoUnknown
DeepInfrafp4$1.30 / $2.601M18 tok/sNoNoConfirmed
DeepInfrafp4$1.30 / $2.601Mnot measuredUnknownUnknownUnknown
DigitalOcean Gradient$1.39 / $2.78262K8 tok/sNoNoConfirmed
Alibaba Cloudfp8$1.42 / $2.831M50 tok/sNoYesunknown periodUnknown
Alibaba Cloud$1.42 / $2.831M53 tok/sNoYesunknown periodUnknown
SiliconFlowfp8$1.50 / $3.131M40 tok/sNoNoConfirmed
Novita AI$1.60 / $3.201Mnot measuredUnknownUnknownUnknown
Venice AI$1.65 / $3.301M43 tok/sNoNoConfirmed
AtlasCloudfp4$1.68 / $3.381M36 tok/sNoYesunknown periodUnknown
Basetenfp4$1.74 / $3.48262K113 tok/sNoNoConfirmed
CoreWeavefp8$1.74 / $3.481M12 tok/sNoNoConfirmed
Together AI$1.74 / $3.48512K55 tok/sNoNoConfirmed
Parasailfp8$1.74 / $3.481M47 tok/sNoNoConfirmed
Fireworks AI$1.74 / $3.481M43 tok/sNoNoConfirmed
Cloudflare Workers AI$1.74 / $3.48393K36 tok/sNoYesunknown periodUnknown

Across the 23 listings we hold: 19 say they do not train on prompts, 1 say they do and 3 do not say. 11 appear in the zero-retention registry we check; the rest are unknown to us rather than confirmed either way.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

Supported
Not supported
Not published
host gave no parameter list
API features per host
ProviderTool callingJSON outputStrict schema
OpenRouter
DeepSeek
Baidufp8
StreamLakefp8
GMICloudfp8
Ionstreamfp4
Novita AIfp8
Waferfp4
DeepInfrafp4
DeepInfrafp4
DigitalOcean Gradient
Alibaba Cloudfp8
Alibaba Cloud
SiliconFlowfp8
Novita AI
Venice AI
AtlasCloudfp4
Basetenfp4
CoreWeavefp8
Together AI
Parasailfp8
Fireworks AI
Cloudflare Workers AI

Tool calling: 20 of 23 listings say yes, 3 publish no parameter list. JSON output: 19 of 23 listings say yes, 1 says no, 3 publish no parameter list. Strict schema: 15 of 23 listings say yes, 5 say no, 3 publish no parameter list.

03

Models people weigh against DeepSeek V4 Pro

04

When we formed this view

Dates behind this page

Aug 3, 2026Price changeStreamLake raised DeepSeek V4 Pro pricing by 29%input +29% ($0.47 → $0.61 per 1M tokens); output +29% ($0.95 → $1.22 per 1M tokens); cache read +29% ($0.039 → $0.051 per 1M tokens)
Aug 2, 2026Price changeDeepSeek V4 Pro cut across 2 hosts, by up to 16% at BaiduDeepSeek V4 Pro moved on 2 hosts: Baidu: input −16% ($0.64 → $0.54 per 1M tokens); output −16% ($1.28 → $1.08 per 1M tokens); cache read −16% ($0.053 → $0.045 per 1M tokens); StreamLake: input −18% ($0.57 → $0.47 per 1M tokens); output −18% ($1.15 → $0.95 per 1M tokens); cache read −18% ($0.048 → $0.039 per 1M tokens)
Aug 2, 2026BenchmarkScored 1502.2 on Arena Codingleaderboard
Aug 2, 2026BenchmarkScored 1446.1 on Arena Creative Writingleaderboard
Aug 2, 2026BenchmarkScored 1480.9 on Arena Hard Promptsleaderboard
Aug 2, 2026BenchmarkScored 1454.7 on Arena Instruction Followingleaderboard
Aug 2, 2026BenchmarkScored 1445.3 on Arena Mathsleaderboard
Aug 2, 2026BenchmarkScored 1457.9 on Arena Text (overall)leaderboard
Aug 2, 2026BenchmarkScored 1446 on Arena Code (WebDev)leaderboard
Aug 1, 2026Price changeDeepSeek V4 Pro repriced across 2 hosts, from a 3% rise at Baidu to a 7% cut at StreamLakeDeepSeek V4 Pro moved on 2 hosts: Baidu: input +3% ($0.63 → $0.64 per 1M tokens); output +3% ($1.25 → $1.28 per 1M tokens); cache read +3% ($0.052 → $0.053 per 1M tokens); StreamLake: input −7% ($0.67 → $0.62 per 1M tokens); output −7% ($1.34 → $1.25 per 1M tokens); cache read −7% ($0.056 → $0.052 per 1M tokens)

Prices last checked 4h ago

What we do not know about this model yet

  • We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
  • 3 of 23 listings publish no parameter list, so what their API accepts is unknown to us.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 3 of 23 listings do not say whether they train on prompts.
05

Licence and identifiers

What the licence allowsMIT License, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.

Licence

MIT License

permissiveCommercial use allowed

Fully permissive: do anything with attribution. No patent grant, unlike Apache-2.0.

Identifiers

Architecture
Mixture of experts
Modality record
text->text
Catalogue slug
deepseek-deepseek-v4-pro

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us