- Type
- Open weightsMIT License
- Params
- 14.7B
- Context
- 16K
about 12K words of context
Our take
Written Aug 2, 2026Microsoft's Phi-4 is a 14.7-billion-parameter text-only model from late 2024 with a permissive MIT licence. It is now clearly behind 2026 peers in both quality and request limit, but it is cheap.
Use this only for legacy pipelines already built on Phi-4 that value stability over quality, or ultra-cheap bulk text processing where a 16,384-token request limit suffices and the quality bar is low. Skip it for new projects or anything requiring vision or a long context.
The case for it
- Permissive MIT licence and low price.
The case against it
- Lowest measured chat quality in our tracked set.
- Tiny request limit by 2026 standards: 16,384 tokens versus 262,144 for current small models.
- Text-only; no vision input, unlike current small-model peers.
How good is it?
IntelligencePuzzles, maths, exam questions
Arena Text (overall)135th of 143 · 1256.2
CodingWriting and fixing code on its own
Arena Coding131st of 143 · 1306.4
Arena Coding is the only board that has scored it for this.
AgenticPlanning, calling tools, staying on task
Nobody we watch has scored Phi 4 for this. We would take the rating from Arena Agent (IPS).
WritingWe do not rate this
Two boards come close and neither tests writing: Arena Creative Writing asks people which of two replies they prefer, and LiveBench Language tests whether a model understood a passage. So we show where Phi 4 placed and give it no mark out of five.
These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.
Every published score for this model9 scoresEvery figure we hold, from 9 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
- Fits in memory
- weights load entirely on the card
- Spills to system RAM
- some weights offload; much slower
- Too large
- will not load even with offload
- est
- size is calculated; the verdict could change by 10%
Comfortable fit
GeForce RTX 4090 · 24 GB
Room to spare. 11.9 GB spare means a 10% error in the size would not change the answer.
GeForce RTX 5090 · 32 GB
Room to spare. 19.9 GB spare means a 10% error in the size would not change the answer.
Apple M1 (8-core GPU) · 16 GB
Room to spare. 1.1 GB spare means a 10% error in the size would not change the answer.
Memory use by level
Against a 24 GB card.
Check against your own machine → · All 71 devices, with every size →
Or rent it from someone else
Cheapest of 3 live listings. Picked at the widest standard context we hold, within one quantisation slice, so the numbers beside it are a price one host actually charges.
- per 1M tokens
- $0.070 in / $0.14 out
- Context served
- 16K
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| OpenRouter | $0.070 / $0.14 | 16K | not measured | Unknown | Unknown | Unknown |
| DeepInfrabfloat16 | $0.070 / $0.14 | 16K | not measured | Unknown | Unknown | Unknown |
| DeepInfrabf16 | $0.070 / $0.14 | 16K | 79 tok/s | No | No | Confirmed |
Across the 3 listings we hold: 1 say they do not train on prompts, 0 say they do and 2 do not say. 1 appear in the zero-retention registry we check; the rest are unknown to us rather than confirmed either way.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
- ✓
- Supported
- ✗
- Not supported
- Not published
- host gave no parameter list
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| OpenRouter | ✗ | ✓ | ✓ |
| DeepInfrabfloat16 | |||
| DeepInfrabf16 | ✗ | ✓ | ✓ |
Tool calling: 0 of 3 listings say yes, 2 say no, 1 publishes no parameter list. JSON output: 2 of 3 listings say yes, 1 publishes no parameter list. Strict schema: 2 of 3 listings say yes, 1 publishes no parameter list.
When we formed this view
Dates behind this page
Prices last checked 3d ago
What we do not know about this model yet
- 1 of 3 listings publish no parameter list, so what their API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 2 of 3 listings do not say whether they train on prompts.
- We hold no cached-input rate for any of its listings.
Licence and identifiers
What the licence allowsMIT License, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
MIT License
Fully permissive: do anything with attribution. No patent grant, unlike Apache-2.0.
Identifiers
- Hugging Face
- microsoft/phi-4
- Architecture
- Dense
- Modality record
- text->text
- Catalogue slug
- microsoft-phi-4