Methodology

What we verify

LLMap is a navigable archive of AI inference: models, hardware, and providers, priced and benchmarked from live sources. This page states what earns a place in the catalogue, how fresh the numbers are, and where our judgment stops.

Listed vs published

Listed means a model, provider, or hardware item appears in the catalogue with factual metadata we ingest from upstream APIs and public pages: names, parameters, prices where available, licence tags, and fit maths. Listing does not imply we recommend it.

Published editorial (the verdict block, what it is for, who should pick it, the catch) appears once automated quality gates pass: numeric grounding against live data, contradiction checks, and denylist rules. Where editorial is not yet ready, the page shows an honest factual line — never a silent gap.

How providers are grouped

The provider list is grouped by what we can show you about each host, strongest evidence first. A band is not a quality ranking and it is not our opinion of a company: each one is a fact you can check on the row itself.

Every row also carries answers that stand on their own, and you can narrow the list to any of them: EU endpoint, No training on prompts, Zero-retention option and SOC 2 attested. The filter is in the address — /providers?eu=yes&zero-retention=yes — so a narrowed list is one you can send to someone else. Narrowing on an answer also removes every host with no answer on record. That is a gap in our records rather than a "no" from them, which is why the unnarrowed list is the one that tells you how much we actually know.

Audited hosts
An independent SOC 2 audit on file with us.
Broad hosts
No audit on file with us, and 20 or more of the models we list on current prices.
Focused hosts
No audit on file with us, and fewer than 20 of the models we list.

Two things follow from that wording, and both are deliberate. A host outside the first band is one we hold no attestation for, which is a gap in our records rather than a failed audit. And a model count counts models in this catalogue, not a provider's own menu, because our view of what they serve arrives through the aggregators. Each row carries the date that host's terms last moved — a price, a context window, a served quantisation, or a model newly listed with them — which is not the same thing as the last time we looked, and does not move when we look and nothing has changed.

Router services such as OpenRouter sit in these bands alongside everyone else. They are a different product, a front door to many hosts rather than a single host, and we hold no field that marks one as a router, so we do not claim to separate them.

Two different dates

When we last read it. Prices, offers and benchmark scores are stamped with the moment we last successfully pulled that value from upstream. Scheduled jobs run between six hours and daily depending on the source, and the ages are genuinely uneven: you will see them per row on a provider page and on a model's Where to run table. "Read 2 days ago" is not an apology; it is the honest age of that row.

When it last changed. That is a different date and we keep it separate, because they answer different questions. A date we publish as "changed" or "last updated" comes from the change itself — a price that moved, a context window that was cut, prose we regenerated — never from the last time a job ran. Otherwise a page would tell you the world had moved when only our reading of it had.

Editorial prose is regenerated when underlying facts change materially. A stale price in prose is a bug; report it.

Fit verdicts and their limits

Fit answers come from measured or estimated weight sizes, your hardware memory, and a throughput model. Three outcomes matter:

  • Fits in memory: the whole model, and the memory it needs to work, sit on the device.
  • Spills to system RAM: part of the model sits in ordinary system memory instead. It may load, and that is not the same as usable — expect far fewer words per second and a longer wait before the first one.
  • Too large: it does not fit even with part of it in system memory.

The weight size is not the whole requirement, which is why a 19.7 GB model can be too large for a 24 GB card. A device keeps some of its memory for the display and the driver, and a loaded model needs room on top of its weights for the software running it and for the conversation it is holding. Where a verdict turns on that difference, the card says so in figures: what the weights come to, what is needed on top, and what the device actually leaves free.

Nearly every verdict is computed from a weight size we calculate from the parameter count rather than read from a published file, and that arithmetic is around 10% out at worst against the files we can check. Where an error that size would change a verdict, we mark it with a question mark, and for a model that runs we say how much spare memory the call turns on. Verdicts with room to spare are not marked, because a warning on everything is a warning on nothing.

A "yes" without the speed consequence is the failure mode we are built to avoid. Benchmark scores and prices are point-in-time facts; fit is a physics estimate, not a guarantee on your exact driver stack.

How to read the marks

The same few marks recur across model, provider, and hardware pages. Each one carries its meaning where it sits — tap or hover a mark and it says what it means in words — so this is a reference, not a key you have to memorise before the pages make sense. No mark relies on colour alone.

✓✗Supported · not supported · not published
A tick means we hold a positive answer — the host supports it. A cross means we hold a negative one. A cross that works against you (say, "trains on your prompts") is outlined as well as coloured, so the direction survives without seeing the colour. A hollow ring is neither: we have no answer on file. It is never a "no" and never a quiet "you're fine" — unknown is unknown, which matters most on privacy.
Rating dots — filled out of five
Each axis (intelligence, coding, writing, agents, transcription) shows how far a model sits behind the leader of one leaderboard, banded onto five dots. The line beneath names the board and the rank — "6th of 144" — because the dots only mean something against that field. They are never a score added up across leaderboards, and never come from a board too small to rank against.
‡via MaxA score measured under a named setting
“via” marks a score measured with a named setting or harness, not with the bare model. “via High”, “via xHigh” and “via Max” name the reasoning-effort setting the leaderboard ran the model at, which decides how much it reasons before it answers; “via mini-SWE-agent” names the harness, the software around the model, that it ran inside. The words are the leaderboard’s own, passed through untranslated, so a score carrying one is not directly comparable with a score without. Where a page has room for the words it prints them; beside a rank line, where it does not, the same thing is drawn as a ‡. Both open this explanation on a tap.
estmeasuredWeight sizes
A size tagged measured was read from a published weights file. One tagged est is our own arithmetic from the parameter count, around 10% out at worst — the tag travels with the number so an estimate never wears measured authority.
A blank price
Where a price cell is empty, the provider does not publish that rate. It is missing data, not free and not zero — we never fill a gap with a number no one charges.

What we redistribute

Where we show a third-party number directly, such as a benchmark score or a provider's list price, we name the source on the page and stamp it with the date we read it. We do not republish upstream datasets, model weights, or raw API responses. Prices and scores are facts about the market, and we present them as such rather than mirroring anyone's data.

Everything else on the site, including how we normalise and reconcile figures that disagree across sources, is our own work. We draw data from the sources below. Each name opens that source's own site rather than the address our pipeline reads, and a name with no link is one we draw on from several places at once, with no single page to send you to:

Report an error

Model, provider, and hardware detail pages include a Report an issue link in the footer area. It opens a prefilled message with the entity slug, page URL, and sync date so we can reproduce the row you saw.

If you would rather just tell us, write to hello@llmap.ai. A wrong number is worth reporting even if you are not certain, and corrections to prices and fit maths get priority.