Providers /Venice AI
Inference provider · HQ US

Venice AI

Models
30
Regions
00

Our take

Editorial by LLMap · updated Aug 3, 2026

Venice AI is a privacy-first US provider that hosts downloadable models with an explicit zero-retention option and no training on prompts. Its catalogue spans 30 open-weight offers with throughput that varies dramatically from one model to another.

Use this for privacy-sensitive workloads where a zero-retention guarantee is required. Pick it for high-throughput inference on select models — CognitiveComputations Uncensored and Xiaomi MiMo V2.5 both exceed 70 tps. Choose it for exploratory access to less-common options such as Z AI GLM 4 7 Flash. Skip it if you need independently verified SOC 2, a disclosed EU-region endpoint, or consistently fast throughput across every model in your stack.

Strengths
  • Explicit zero-retention option for privacy-conscious users, plus a commitment not to train on prompts.
  • Strong peak throughput on select models — up to 84 tps on CognitiveComputations Uncensored and 71 tps on Xiaomi MiMo V2.5, which is 6.7× the throughput of its slowest tracked model.
  • Low entry price on some models in the catalogue.
Trade-offs
  • Compliance and regional infrastructure gaps: SOC 2 is unverified in our data and no EU-region endpoint is disclosed.
  • Inconsistent throughput across the catalogue — DeepSeek V3.2 and Mistral Small 3.2 24B both fall below 20 tps while top models exceed 70 tps.
  • Some models are priced at a premium versus cheaper alternatives in the same catalogue: DeepSeek V3.2 costs 5.5× the input price of Z AI GLM 4 7 Flash and 3.2× the output price of Qwen Qwen3 5 9B.
01

Privacy & data handling

Yes
No
No answer on record

The colour says whether the answer favours you, not whether it is a yes — not training on your prompts earns a green cross, no zero-retention option earns a red one.

Trains on your prompts
No
Logs prompts
No
Offered
No answer on record

Prompt retention: none.

02

Models & pricing

Something wrong on this page? Tell us