Models / Best of /cheapest
Updated October 2026 · ranked from live data

Cheapest hosted models

Ranked by the cheapest like-for-like listing we hold for each model — standard tier, at the widest context that model is served at, within one quantisation — one real offer per row with the host named. A model page may show a cheaper row that is quantised or serves less context. Price is the only axis here: a row near the top is cheap, not good. Prices sync daily. Right now the top pick is Llama 3.2 1B Instruct ($0.020/M out). Every row links to full specs, hardware verdicts and provider pricing.

1Llama 3.2 1B InstructMeta$0.020/M out1.2B$0.020 / $0.020per 1M · via Novita AI2Ling-2.6-flashinclusionAI$0.030/M out107B$0.010 / $0.030per 1M · via OpenRouter3Mistral NemoMistral AI$0.030/M out12.2B$0.019 / $0.030per 1M · via OpenRouter4Llama 3 8B LunarisSao10K$0.050/M out8B$0.040 / $0.050per 1M · via OpenRouter5Ling-3.0-flashinclusionAI$0.062/M out128B$0.021 / $0.062per 1M · via OpenRouter6Mistral Small 3Mistral AI$0.080/M out23.6B$0.050 / $0.080per 1M · via OpenRouter7Llama 3.1 8B InstructMeta$0.080/M out8B$0.050 / $0.080per 1M · via Groq8DeepSeek V4 FlashDeepSeek$0.084/M out291B / 37B$0.042 / $0.084per 1M · via OpenRouter9Nex-N2.5-MiniNex AGI$0.10/M out35.1B$0.025 / $0.10per 1M · via OpenRouter10Granite 4.1 8BIBM$0.10/M out8.8B$0.050 / $0.10per 1M · via OpenRouter
More listsBest models for codingBest models for writingBest models that fit in 12 GB VRAMBest models for 24–32 GB VRAMLongest context windowsBest open-weight modelsFastest hosted models