Hardware / NVIDIA /RTX 6000 Ada
NVIDIA · GPU
RTX 6000 Ada
Runs 116 models fully in memory, 14 more with CPU offload.
Memory
48 GB
Bandwidth
960 GB/s
Runs
116
01
Specifications
Memory
48 GB
Bandwidth
960 GB/s
FP16 compute
91.1 TFLOPS
TDP
300 W
MSRP
$6,800
Street price
—
Released
Dec 3, 2022
Architecture
Ada Lovelace
02
What it runs
Granite 4.1 8B8.8BQ4_K_M131K144 tok/sestFits in memoryQwen3.6 35B A3B36BQ4_K_M262K359 tok/sestFits in memoryGemma 4 26B A4B26.5BQ5_K_M66K230 tok/sestFits in memoryQwen3.5-9B9.7BQ4_K_M262K131 tok/sestFits in memoryNemotron 3 Nano 30B A3B31.6BQ5_K_M262K306 tok/sestFits in memoryMinistral 3 8B 25128.9BQ4_K_M262K142 tok/sestFits in memoryMinistral 3 3B 25123.8BQ5_K_M131K284 tok/sestFits in memoryGranite 4.0 Micro3.2BQ8_066K192 tok/sestFits in memoryQwen3 VL 8B Thinking8.8BQ5_K_M131K123 tok/sestFits in memoryQwen3 VL 30B A3B Thinking31.1BQ5_K_M131K306 tok/sestFits in memoryUI-TARS 7B8.3BQ5_K_M66K130 tok/sestFits in memoryLlama 3.2 3B Instruct3.2BQ5_K_M131K337 tok/sestFits in memoryQwen3.5-35B-A3B36BQ4_K_M262K359 tok/sestFits in memoryQwen3 VL 8B Instruct8.8BQ5_K_M262K123 tok/sestFits in memoryQwen3 VL 30B A3B Instruct31.1BQ4_K_M131K359 tok/sestFits in memoryQwen3 30B A3B Thinking 250730.5BQ4_K_M66K359 tok/sestFits in memoryQwen3 Coder 30B A3B Instruct30.5BQ5_K_M131K306 tok/sestFits in memoryQwen3 30B A3B Instruct 250730.5BQ4_K_M262K359 tok/sestFits in memoryQwen3 8B8.2BQ5_K_M131K132 tok/sestFits in memoryGemma 3 4B4.3BQ8_0131K168 tok/sestFits in memoryLlama 3.1 8B Instruct8BQ5_K_M131K135 tok/sestFits in memoryNiagara 19m Batch.en20MQ8_0~262K36141 tok/sestFits in memoryNiagara 38m Batch.en38MQ5_K_M~262K28416 tok/sestFits in memoryARK ASR 0.6B1.3BQ8_0~262K556 tok/sestFits in memoryARK ASR 3B4.1BQ8_0~262K176 tok/sestFits in memoryAudio8 ASR 0.1B0.3BQ8_0~262K2409 tok/sestFits in memoryHiggs Audio v3 8b STT v28.9BQ4_K_M~262K142 tok/sestFits in memoryHiggs Audio v3 STT2.7BQ5_K_M~262K400 tok/sestFits in memoryCohere Transcribe 03 20262.1BQ5_K_M~262K514 tok/sestFits in memoryDistil Large v3.50.8BQ4_K_M~262K1584 tok/sestFits in memoryLite Whisper Large v3 Acc1.4BQ8_0~262K516 tok/sestFits in memoryOwsm CTC v3.1 1B1.1BQ8_0~262K657 tok/sestFits in memoryOwsm CTC v3.2 ft 1B1BQ4_K_M~262K1267 tok/sestFits in memoryOwsm CTC v4 1B1BQ4_K_M~262K1267 tok/sestFits in memoryHubert Large Ls960 ft0.3BQ4_K_M~262K3959 tok/sestFits in memoryMMS 1b All1BQ8_0~262K723 tok/sestFits in memoryWav2vec2 Large 960h Lv60 Self0.3BQ8_0~262K2259 tok/sestFits in memoryHojo ASR V15.2BQ8_0~262K140 tok/sestFits in memoryGranite 4.0 1b Speech2.3BQ4_K_M~262K551 tok/sestFits in memoryGranite Speech 3.3 2b3BQ5_K_M~262K360 tok/sestFits in memoryGranite Speech 3.3 8b8.6BQ4_K_M~131K147 tok/sestFits in memoryGranite Speech 4.1 2b2.3BQ4_K_M~262K551 tok/sestFits in memoryGranite Speech 4.1 2b NAR2.3BQ5_K_M~262K470 tok/sestFits in memorySTT 2.6b en2.6BQ5_K_M~262K415 tok/sestFits in memoryPhi 4 Multimodal Instruct5.6BQ4_K_M~262K226 tok/sestFits in memoryVibeVoice ASR HF8.3BQ4_K_M~262K153 tok/sestFits in memoryVoxtral Mini 3B 25075BQ4_K_M~262K253 tok/sestFits in memoryVoxtral Mini 4B Realtime 26024.4BQ4_K_M~262K288 tok/sestFits in memoryCanary 180m Flash0.2BQ4_K_M~262K7038 tok/sestFits in memoryCanary 1b1BQ5_K_M~262K1080 tok/sestFits in memoryCanary 1b Flash1BQ4_K_M~262K1267 tok/sestFits in memoryCanary 1b v21BQ4_K_M~262K1267 tok/sestFits in memoryCanary Qwen 2.5b2.6BQ5_K_M~262K415 tok/sestFits in memoryNemotron 3.5 ASR Streaming 0.6b0.6BQ5_K_M~262K1687 tok/sestFits in memoryNemotron Speech Streaming en 0.6b0.6BQ8_0~262K1205 tok/sestFits in memoryParakeet CTC 0.6b0.6BQ4_K_M~262K2111 tok/sestFits in memoryParakeet CTC 1.1b1.1BQ5_K_M~262K982 tok/sestFits in memoryParakeet RNNT 0.6b0.6BQ8_0~262K1205 tok/sestFits in memoryParakeet RNNT 1.1b1.1BQ4_K_M~262K1152 tok/sestFits in memoryParakeet TDT CTC 110m0.1BQ8_0~262K6571 tok/sestFits in memory
A speed marked est is calculated from this card's memory bandwidth and the size of the weights it has to read per word. It is not a measurement.