Our take
OpenInference is a small hosted inference provider with a narrow catalogue of just four offers across two model families. It stands out for a zero-retention data option and a commitment not to train on prompts, though compliance documentation remains sparse.
Use this when zero data retention is a hard requirement, or for low-volume inference on DeepSeek V4 Flash where the cheapest tracked tier here is the priority. Pick it for Gemma 4 31B access at moderate cost with the highest throughput in this catalogue. Skip it if you need a broad model choice, verified SOC 2, a confirmed EU endpoint, or predictable throughput that does not degrade on cheaper tiers.
- Cheapest tracked DeepSeek V4 Flash tier in this catalogue.
- Zero-retention data option available, with no training on prompts.
- Gemma 4 31B served at 31 tokens per second, the highest throughput here.
- Very narrow catalogue: only four tracked offers across two model families.
- Throughput degrades sharply on cheaper DeepSeek tiers — the most expensive tier is roughly half the speed of the cheapest despite costing more.
- Compliance and geographic transparency gaps: headquarters, EU endpoint and SOC 2 all unverified in our data.
Point your tools here
We have not recorded what a router needs for OpenInference yet — no base URL and no docs link on file. That is our gap, not a sign OpenInference has no API; its own documentation is the place to look until we close it.
Privacy & data handling
These answers cover requests sent to OpenInference through OpenRouter, as OpenRouter records them, last read 4 hours ago. For OpenInference's own API we hold no answer: check its terms. SOC 2 and the data processing agreement below are hand-checked, not part of that sync.
Prompt retention: none.