Our take
A boutique GPU-cloud inference host with a tight, curated catalogue of seven models. Most offers include measured throughput in tokens per second, making it useful when you need verified speed data rather than a large model selection.
Use this for throughput-sensitive workloads where verified tokens-per-second matters: six of seven offers have measured throughput. Pick it if you prefer the NVIDIA ecosystem, as the Nemotron Nano anchor offers 80 tps at $0.05/$0.20 per million tokens. Skip it if you need broad model coverage or low prices on frontier-class models.
- Strong measured throughput on most models — six of seven tracked offers have throughput data, with z-ai-glm-5-1 at 82 tps and nvidia-nemotron-3-nano-30b-a3b at 80 tps.
- Clear price/performance anchor — nvidia-nemotron-3-nano-30b-a3b at $0.05/$0.20 with 80 tps is the cheapest and fastest tracked offer.
- Very small catalogue limits model coverage — only seven tracked offers versus 61 at volume leaders.
- Premium pricing on frontier-class models: moonshotai-kimi-k2-6 at $0.70/$3.50 and z-ai-glm-5-1 at $1.20/$4.40 are among the highest prices tracked for these models.
- DeepSeek V3 throughput lags within the same fleet at 32 tps, well below the 80–82 tps seen on other offers.
Privacy & data handling
- ✓
- Yes
- ✗
- No
- No answer on record
The colour says whether the answer favours you, not whether it is a yes — not training on your prompts earns a green cross, no zero-retention option earns a red one.
Prompt retention: none.