Methodology
What we verify
LLMap is a navigable archive of AI inference: models, hardware, and providers, priced and benchmarked from live sources. This page states what earns a place in the catalogue, how fresh the numbers are, and where our judgment stops.
Listed vs published
Listed means a model, provider, or hardware item appears in the catalogue with factual metadata we ingest from upstream APIs and public pages: names, parameters, prices where available, licence tags, and fit maths. Listing does not imply we recommend it.
Published editorial (the verdict block, what it is for, who should pick it, the catch) appears once automated quality gates pass: numeric grounding against live data, contradiction checks, and denylist rules. Where editorial is not yet ready, the page shows an honest factual line — never a silent gap.
Provider tiers
We do not publish a caste system at launch. Provider rows expose observable filters you can trust: EU endpoint, no-training policy, zero-retention option, SOC 2 where verified. Router services such as OpenRouter are grouped separately because they are a different product: a front door to many models rather than a single host.
A published Established and Emerging split will arrive only once we can state the criteria behind it and show when each provider was last checked. Until then, treat tier labels as work in progress rather than a ranking.
Freshness and "synced"
Every price, offer, and benchmark score carries a synced timestamp: when we last successfully ingested that value from upstream. Scheduled jobs run between six hours and daily depending on the source. "Synced 2 days ago" is not an apology. It is the honest age of that row.
Editorial prose is regenerated when underlying facts change materially. A stale price in prose is a bug; report it.
Fit verdicts and their limits
Fit answers come from measured or estimated weight sizes, your hardware memory, and a throughput model. Three outcomes matter:
- Runs: the quant fits fully in VRAM at usable speed.
- Offloads: layers spill to system RAM. The model may load, and that is not the same as usable. Expect sharply lower tokens per second and higher latency. We surface an explicit spill warning and an estimated speed where we have one.
- Won't fit: even with offload, the configuration is not practical.
Nearly every verdict is computed from a weight size we calculate from the parameter count rather than read from a published file, and that arithmetic is around 10% out at worst against the files we can check. Where an error that size would change a verdict, we mark it with a question mark, and for a model that runs we say how much spare memory the call turns on. Verdicts with room to spare are not marked, because a warning on everything is a warning on nothing.
A "yes" without the speed consequence is the failure mode we are built to avoid. Benchmark scores and prices are point-in-time facts; fit is a physics estimate, not a guarantee on your exact driver stack.
What we redistribute
Where we show a third-party number directly, such as a benchmark score or a provider's list price, we name the source on the page and stamp it with the date we read it. We do not republish upstream datasets, model weights, or raw API responses. Prices and scores are facts about the market, and we present them as such rather than mirroring anyone's data.
Everything else on the site, including how we normalise and reconcile figures that disagree across sources, is our own work. We draw data from:
Report an error
Model, provider, and hardware detail pages include a Report an issue link in the footer area. It opens a prefilled message with the entity slug, page URL, and sync date so we can reproduce the row you saw.
If you would rather just tell us, write tohello@llmap.ai. A wrong number is worth reporting even if you are not certain, and corrections to prices and fit maths get priority.