Our take
A small inference host with competitive pricing on DeepSeek models and a cluster of Z-AI GLM family offerings. It offers the highest throughput in its catalogue on the entry-level DeepSeek V4 Flash at 69 tps, but has thin documentation on compliance and geographic footprint.
Use this for DeepSeek V4 Flash inference where 69 tps throughput and $0.091/$0.182 per million tokens is competitive. Pick it for access to the Z-AI GLM model family, which is not widely hosted elsewhere. Skip it if you need verified SOC 2, a known headquarters country, EU data residency, or a deep catalogue — it has only 7 tracked offers and its compliance posture is entirely undocumented.
- Highest throughput in catalogue on entry-level DeepSeek model: DeepSeek V4 Flash at 69 tps, above V4 Pro at 51 tps, V3-2 at 43 tps, and all other tracked offers.
- Competitive pricing on DeepSeek V4 Flash at $0.091/$0.182 per million tokens — about one-seventh the price of V4 Pro at $0.6253/$1.2506.
- Very small catalogue depth: 7 tracked offers versus 61 at catalogue leaders.
- Throughput drops sharply on premium models: V4 Pro at 51 tps, V3-2 at 43 tps, Kimi K2-6 at 39 tps, and GLM-5-1 at 37 tps — all below V4 Flash's 69 tps.
- Compliance and data residency posture is entirely undocumented: headquarters country, EU endpoint, SOC 2, training-on-prompts policy, and zero-retention are all unverified.
Privacy & data handling
- ✓
- Yes
- ✗
- No
- No answer on record
The colour says whether the answer favours you, not whether it is a yes — not training on your prompts earns a green cross, no zero-retention option earns a red one.