Our take
A boutique inference host with a tight, Asia-weighted model catalogue and measured throughput data for every offer. It is useful when you need speed benchmarks to compare hosts, but thin on compliance and geographic disclosures.
Use this when throughput is a deciding factor and you need measured tokens-per-second figures to compare hosts. Pick it for access to specific Asian lab models — MiniMax M2.5, Z-AI GLM 5, and Qwen3 235B-A22B — that larger catalogues may not carry. Skip it if you need a broad model selection, independently verified compliance, or documented data-sovereignty guarantees.
- Provides throughput benchmarks for every tracked offer — all 6 offers include measured tps, with MiniMax M2.5 at 89 tps and Z-AI GLM 5-2 at 88 tps as the fastest.
- Carries models from Asian labs not universally hosted: MiniMax M2.5, Z-AI GLM 5, and Qwen3 235B-A22B in a 6-offer catalogue.
- Very small catalogue versus open-weight hosts — 6 tracked offers, compared with 61 at catalogue leaders.
- Premium pricing on several offers: Z-AI GLM 5-1 and 5-2 at $1.40/$4.40 per million tokens, and DeepSeek V3-2 at $0.50/$1.50, with lower price floors elsewhere for comparable models.
- Compliance and data-sovereignty posture is entirely undocumented: HQ country, EU endpoint, SOC 2, zero retention, and training-on-prompts are all unverified in our data.
Privacy & data handling
- ✓
- Yes
- ✗
- No
- No answer on record
The colour says whether the answer favours you, not whether it is a yes — not training on your prompts earns a green cross, no zero-retention option earns a red one.