Our take
CoreWeave is a GPU-cloud inference provider that publishes measured throughput on every tracked offer, with a zero-retention data option for privacy-sensitive workloads. It carries 25 offers but leaves compliance and geographic details unverified in our data.
Use this when throughput is a decision factor — every offer includes measured tokens per second. Pick it when data retention control matters, or for budget inference on smaller models. Skip it if you need verified SOC 2, a confirmed EU endpoint, or a known headquarters jurisdiction.
- Transparent throughput data on every tracked offer — all 12 sample offers include measured tps, ranging from 28 to 129.
- Strong throughput on select models: nvidia-nemotron-3-5-lightning at 129 tps, deepseek-deepseek-v4-flash at 118 tps.
- Zero-retention option available; does not train on prompts.
- Competitive pricing on small-parameter models.
- Compliance and geographic documentation is thin: SOC 2, EU endpoint and headquarters country are all unverified in our data.
- Throughput inconsistent on the same model: google-gemma-4-31b shows 28 tps and 58 tps, a more than 2x spread between two readings.
- Some models show low throughput: google-gemma-4-31b at 28 tps, qwen-qwen3-30b-a3b-instruct-2507 at 29 tps.
Point your tools here
We have not recorded what a router needs for CoreWeave yet — no base URL and no docs link on file. That is our gap, not a sign CoreWeave has no API; its own documentation is the place to look until we close it.
Privacy & data handling
These answers cover requests sent to CoreWeave through OpenRouter, as OpenRouter records them, last read 3 hours ago. For CoreWeave's own API we hold no answer: check its terms. SOC 2 and the data processing agreement below are hand-checked, not part of that sync.
Prompt retention: none.