Models / Mistral AI/ Mistral Nemo

Mistral Nemo

Mistral AI · released Jul 17, 2024 · mistralai/Mistral-Nemo-Instruct-2407

Input: text. Output: text.InputOutput
Type
Open weightsApache License 2.0
Params
12.2B
Context
131K

about 98K words of context

Our take

Written Aug 3, 2026

Mistral Nemo is a compact downloadable text model with a permissive Apache licence and a generous request limit for its size. Its academic scores are near the floor for modern models, so it is a budget utility pick rather than a quality leader.

Who should pick it

Pick this for low-cost text inference where Apache licensing matters, or for long-context text tasks up to 131,072 tokens. Use it for local or self-hosted deployment with open weights. Skip it if you need strong reasoning or knowledge accuracy, or if you want measured chat or code quality.

The case for it

  • Among the cheapest hosted inference we track at the low end.
  • Apache 2.0 licence allows commercial use, fine-tuning and redistribution.
  • 131,072-token request limit is large for a 12.2-billion-parameter model.

The case against it

  • Weak academic benchmark performance: 28% on MMLU-Pro and 5.4% on GPQA Diamond, both near floor levels for modern models.
  • Wide price spread with no throughput advantage at premium tiers: the vendor's own API charges several times more than the cheapest hosts for 48 tps, while DeepInfra offers 34 tps at the budget rate.
00

How good is it?

IntelligencePuzzles, maths, exam questions

Scored, not ratedGPQA Diamond · 12th of 16 · 5.4

Mistral Nemo is not on Arena Text (overall), which is where the rating would come from, so there is no rating here. It is on GPQA Diamond, in 12th of 16 with 5.4.

MMLU-Pro 13th of 16

CodingWriting and fixing code on its own

not measured

Nobody we watch has scored Mistral Nemo for this. We would take the rating from Arena Coding.

AgenticPlanning, calling tools, staying on task

not measured

Nobody we watch has scored Mistral Nemo for this. We would take the rating from Arena Agent (IPS).

WritingWe do not rate this

not measured

Nobody we watch has scored this model for writing. Two boards come close and neither tests writing: Arena Creative Writing asks people which of two replies they prefer, and LiveBench Language tests whether a model understood a passage.

Also scored, on boards we give no mark for
IFEval 12th of 16

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.

Every published score for this model3 scoresEvery figure we hold, from 3 boards, with who ran it and a link to the source — including the boards no rating above is built on.
GPQA Diamondreasoning
5.4independentsource ↗
IFEvalchat
63.8independentsource ↗
MMLU-Proreasoning
28independentsource ↗
01

Can you run it yourself?

Fits in memory
weights load entirely on the card
Spills to system RAM
some weights offload; much slower
Too large
will not load even with offload
est
size is calculated; the verdict could change by 10%

Comfortable fit

A card many people ownFits in memory

GeForce RTX 4090 · 24 GB

Weights at Q4_K_M7.7 / 24 GBest
Spare memory13.3 GB spare
Usable context66K of 131K
Decode speed109 tok/sest

Room to spare. 13.3 GB spare means a 10% error in the size would not change the answer.

One step upFits in memory

GeForce RTX 5090 · 32 GB

Weights at Q4_K_M7.7 / 32 GBest
Spare memory21.3 GB spare
Usable context131K of 131K
Decode speed194 tok/sest

Room to spare. 21.3 GB spare means a 10% error in the size would not change the answer.

On a MacFits in memory

Apple M1 (8-core GPU) · 16 GB

Weights at Q4_K_M7.7 / 16 GBest
Spare memory2.5 GB spare
Usable context16K of 131K
Decode speed6 tok/sest

Room to spare. 2.5 GB spare means a 10% error in the size would not change the answer.

Your hardware
Checking your profile…

Memory use by level

Against a 24 GB card.

Q4_K_M
recommended
7.7 GBest
Fits in memory
Q5_K_M
9 GBest
Fits in memory
Q8_0
13.5 GBest
Fits in memory

All 3 sizes here are calculated, not measured. We hold no measured file for this model, so each size comes from the parameter count and every verdict above inherits that uncertainty.

Check against your own machine → · All 71 devices, with every size →

02

Or rent it from someone else

Cheapest published offer

Cheapest of 8 live listings. Picked at the widest standard context we hold, within one quantisation slice, so the numbers beside it are a price one host actually charges.

per 1M tokens
$0.019 in / $0.030 out
Context served
131K
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
DeepInfrafp8$0.019 / $0.030131Knot measuredUnknownUnknownUnknown
DeepInfrafp8$0.019 / $0.030131K34 tok/sNoNoConfirmed
OpenRouter$0.019 / $0.030131Knot measuredUnknownUnknownUnknown
Parasailfp8$0.030 / $0.030131K20 tok/sNoNoConfirmed
Mistral AI$0.15 / $0.15131K48 tok/sNoYes30 daysUnknown
Io Netfp16$0.043 / $0.17128K9 tok/sNoNoConfirmed
Novita AIfp8$0.040 / $0.1760K21 tok/sNoNoConfirmed
Novita AI$0.040 / $0.1760Knot measuredUnknownUnknownUnknown

Across the 8 listings we hold: 5 say they do not train on prompts, 0 say they do and 3 do not say. 4 appear in the zero-retention registry we check; the rest are unknown to us rather than confirmed either way.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

Supported
Not supported
Not published
host gave no parameter list
API features per host
ProviderTool callingJSON outputStrict schema
DeepInfrafp8
DeepInfrafp8
OpenRouter
Parasailfp8
Mistral AI
Io Netfp16
Novita AIfp8
Novita AI

Tool calling: 3 of 8 listings say yes, 3 say no, 2 publish no parameter list. JSON output: 5 of 8 listings say yes, 1 says no, 2 publish no parameter list. Strict schema: 4 of 8 listings say yes, 2 say no, 2 publish no parameter list.

03

Models people weigh against Mistral Nemo

04

When we formed this view

Dates behind this page

Aug 3, 2026BenchmarkScored 5.4 on GPQA Diamondleaderboard
Aug 3, 2026BenchmarkScored 63.8 on IFEvalleaderboard
Aug 3, 2026BenchmarkScored 28 on MMLU-Proleaderboard
Jul 31, 2026Price changeIo Net raised Mistral Nemo pricing by 10%input +10% ($0.039 → $0.043 per 1M tokens); output +10% ($0.15 → $0.17 per 1M tokens); cache read +10% ($0.019 → $0.021 per 1M tokens)
Jul 30, 2026Price changeIo Net raised Mistral Nemo pricing by 15%input +11% ($0.035 → $0.039 per 1M tokens); output +15% ($0.13 → $0.15 per 1M tokens); cache read +11% ($0.018 → $0.019 per 1M tokens)
Jul 26, 2026ListedListed on LLMapfirst indexed by our pipeline
Jul 17, 2024AnnouncedMistral Nemo announced by Mistral AI

Prices last checked 5d ago

What we do not know about this model yet

  • We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
  • 2 of 8 listings publish no parameter list, so what their API accepts is unknown to us.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 3 of 8 listings do not say whether they train on prompts.
05

Licence and identifiers

What the licence allowsApache License 2.0, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.

Licence

Apache License 2.0

permissiveCommercial use allowed

Fully permissive: commercial use, redistribution, and derivatives allowed. Requires attribution and a copy of the license. Includes an express patent grant.

Identifiers

Architecture
Dense
Modality record
text->text
Catalogue slug
mistralai-mistral-nemo

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us