Models / Best of /fastest
Updated October 2026 · ranked from live data

Fastest hosted models

General-purpose text chat models only, ranked by the median throughput across the hosts serving them. The figures are OpenRouter’s own measurement of each endpoint over the last 30 minutes, not a benchmark we ran. Specialist and agentic routes are excluded; a model needs at least two separate hosts measured, so one outlier cannot decide the order. Right now the top pick is Nemotron 3.5 Lightning (155 tok/s median · 5 hosts). Every row links to full specs, hardware verdicts and provider pricing.

1Nemotron 3.5 LightningNVIDIA155 tok/s median · 5 hosts31.6B$0.059 / $0.17per 1M · via Io Net2GPT-5.6 Luna ProOpenAI142 tok/s median · 2 hosts—$0.20 / $1.20per 1M · via OpenAI3Qwen3 Next 80B A3B ThinkingQwen139 tok/s median · 2 hosts81.3B / 3B$0.15 / $1.20per 1M · via Google Vertex AI4gpt-oss-120bOpenAI118 tok/s median · 19 hosts120B / 5.1B$0.090 / $0.36per 1M · via Google Vertex AI5Nemotron 3 Nano 30B A3BNVIDIA117 tok/s median · 4 hosts31.6B / 3B$0.050 / $0.20per 1M · via OpenRouter6Gemini 3.8 FlashGoogle115 tok/s median · 2 hosts—$0.75 / $3.75per 1M · via Google AI7Gemini 2.5 Flash LiteGoogle108 tok/s median · 2 hosts—$0.10 / $0.40per 1M · via Google AI8Qwen3 VL 30B A3B ThinkingQwen98 tok/s median · 2 hosts31.1B / 3B$0.20 / $2.40per 1M · via OpenRouter9Gemini 3.7 FlashGoogle97 tok/s median · 2 hosts—$0.75 / $3.75per 1M · via Google AI10GPT-6 Luna ProOpenAI95 tok/s median · 2 hosts—$0.10 / $0.50per 1M · via OpenAI
More listsBest models for codingBest models for writingCheapest hosted modelsBest models that fit in 12 GB VRAMBest models for 24–32 GB VRAMLongest context windowsBest open-weight models