Models / Best of /local-12gb
Updated October 2026 · newest first, from live data

Best models that fit in 12 GB VRAM

Open-weight models that fit a 12 GB card (RTX 3060/4070 class) at Q4 quantisation with room for context, newest first. Newest first, not a ranking by quality. 2 of the 10 below carry a rating from a board we track, and the highest placed is Granite 4.1 8B (1/5 everyday (Arena Text (overall), 145 of 168)). For the other 8 we hold no score at all, which is not the same as a low one. Every row links to full specs, hardware verdicts and provider pricing.

What is quantisation? →
1Granite 4.2 8BIBM1/5 everyday (Arena Text (overall), 153 of 168)est8.8B$0.060 / $0.25per 1M · via OpenRouter2Nemotron 3.5 Content SafetyNVIDIAno rating from any board we trackest4.3B$0.20 / $0.20per 1M · via OpenRouter3Hy-MT2-7BTencentno rating from any board we trackest8B$0.074 / $0.29per 1M · via OpenRouter4Hy-MT2-1.8BTencentno rating from any board we trackest2B$0.044 / $0.18per 1M · via OpenRouter5Granite 4.1 8BIBM1/5 everyday (Arena Text (overall), 145 of 168)est8.8B$0.050 / $0.10per 1M · via OpenRouter6Reka EdgeReka AIno rating from any board we trackest7.1B$0.10 / $0.10per 1M · via Reka7Qwen3.5-9BQwenno rating from any board we trackest9.7B$0.10 / $0.15per 1M · via OpenRouter8Ministral 3 3B 2512Mistral AIno rating from any board we trackest3.8B$0.10 / $0.10per 1M · via Mistral AI9Ministral 3 8B 2512Mistral AIno rating from any board we trackest8.9B$0.15 / $0.15per 1M · via Mistral AI10Ministral 3 14B 2512Mistral AIno rating from any board we trackest13.9B$0.20 / $0.20per 1M · via Mistral AI
More listsBest models for codingBest models for writingCheapest hosted modelsBest models for 24–32 GB VRAMLongest context windowsBest open-weight modelsFastest hosted models