Catalogue
Models
376 models, cross-linked to hardware verdicts and live provider pricing.
78 models
Whisper3 sizes here
Whisper 1 APIOpenAISpeech to textWhisper 1 API is OpenAI's hosted speech-to-text service, available from both OpenAI and Azure at the same per-minute rate.————$0.006/minute of audiovia Microsoft Azure AIWhisper Large v3 TurboOpenAISpeech to textWhisper Large v3 Turbo is a compact downloadable speech-to-text model from OpenAI that turns audio into written words at exceptional speed.93%783×990.8B$0.001/minute of audiovia GroqWhisper Large v3OpenAISpeech to textWhisper Large v3 turns recorded speech into written words, and is still the name most people reach for.93.4%462×991.5B$0.002/minute of audiovia GroqScribe v1ZoomSpeech to textScribe v1 is Zoom's proprietary speech-to-text model, available only through ElevenLabs, that turns English audio into written text.95.3%—1—$0.004/minute of audiovia ElevenLabsSpeechmatics EnhancedSpeechmaticsSpeech to textSpeechmatics Enhanced is a hosted-only speech-to-text engine that turns audio into written words across 55 languages.94.1%—55——no live hostSmallest AI PulseSmallest AISpeech to textSmallest AI Pulse is a hosted-only speech-to-text model that turns audio into written words across 38 claimed languages.94.6%—38——no live hostNemotron 3.5 ASR Streaming 0.6bNVIDIASpeech to textNemotron 3.5 ASR Streaming is a tiny downloadable speech-to-text model from NVIDIA that turns audio into written words across 35 languages.92.1%1,490×350.6B—no live hostModulate VfastModulateSpeech to textModulate Vfast is a hosted-only speech-to-text model that turns audio into written words.95.6%————no live hostAzure Speech 06 2026MicrosoftSpeech to textAzure Speech is Microsoft's hosted-only transcription service that turns audio into text across 25 languages.95.5%—25——no live hostSolaria 3GladiaSpeech to textSolaria 3 is Gladia's hosted-only speech-to-text model that turns audio into written words.94.7%—5——no live hostomniASR LLM 7B v2MetaSpeech to text93.1%143×16767.8B—no live hostomniASR CTC 7B v2MetaSpeech to text91%528×16766.5B—no live hostScribe v2ElevenLabsSpeech to textScribe v2 is ElevenLabs' proprietary speech-to-text model that turns audio into written words across 90 languages.95.4%—90——no live hostAvalon v1 enAqua VoiceSpeech to textAvalon v1 is Aqua Voice's proprietary English-only speech-to-text model that excels at clean, prepared audio but falls apart on challenging real-world recordings.94.8%————no live hostQwen3 ASR4 sizes here
Qwen3 ASR 1.7B HFQwenSpeech to textQwen3 ASR is a compact downloadable speech-to-text model with an Apache licence and support for 30 languages.95%796×302B—no live hostQwen3 ASR 1.7BQwenSpeech to textQwen3 ASR is a tiny downloadable speech-to-text model from Qwen that turns audio into written words.95%394×522B—no live hostQwen3 ASR 0.6BQwenSpeech to textQwen3 ASR is a tiny downloadable speech-to-text model with an Apache licence.94.4%439×520.8B—no live hostQwen3 ASR 0.6B HFQwenSpeech to textQwen3 ASR is a tiny downloadable speech-to-text model with an Apache licence and support for 30 languages.94.4%730×300.8B—no live hostParakeet8 sizes here
Parakeet RNNT 1.1bNVIDIASpeech to textParakeet RNNT is a small downloadable speech-to-text model from NVIDIA that turns English audio into written text at extraordinary speed.93.6%4,122×11.1B—no live hostParakeet TDT 0.6b v2NVIDIASpeech to textParakeet is NVIDIA's tiny downloadable speech-to-text model, built to turn English audio into written text at extraordinary speed on modest hardware.94.6%6,038×10.6B—no live hostParakeet RNNT 0.6bNVIDIASpeech to textParakeet RNNT is a tiny downloadable speech-to-text model from NVIDIA that processes audio faster than almost anything else we track.93.4%5,407×10.6B—no live hostParakeet CTC 0.6bNVIDIASpeech to textParakeet CTC is a tiny downloadable speech-to-text model from NVIDIA that turns English audio into written text almost instantly.93.3%5,884×10.6B—no live hostParakeet TDT 0.6b v3NVIDIASpeech to textParakeet TDT is a tiny downloadable speech-to-text model from NVIDIA that turns audio into written words at extraordinary speed — an hour of audio in under a second on benchmark hardware. It is extremely accurate on clean read-aloud content but far less so on accented speech or recorded meetings, and it must be self-hosted as no commercial providers currently offer it.94.3%6,098×250.6B—no live hostParakeet TDT CTC 110mNVIDIASpeech to textParakeet is NVIDIA's tiny downloadable speech-to-text model built for blistering speed on clean English audio.93.4%6,119×10.1B—no live hostParakeet TDT 1.1bNVIDIASpeech to textParakeet TDT 1.1b is NVIDIA's tiny downloadable speech-to-text model that turns English audio into written words at extraordinary speed.93.8%4,529×11.1B—no live hostParakeet CTC 1.1bNVIDIASpeech to textParakeet CTC is a tiny downloadable speech-to-text model from NVIDIA that turns English audio into written words at extraordinary speed.93.5%5,015×11.1B—no live hostVoxtral2 sizes here
Voxtral Mini 3B 2507Mistral AISpeech to textVoxtral Mini is a 5-billion-parameter speech-to-text model from Mistral AI with an Apache licence.94%180×85B—no live hostVoxtral Mini 4B Realtime 2602Mistral AISpeech to textVoxtral Mini is a compact downloadable speech-to-text model from Mistral AI that turns audio into written words.93.6%105×134.4B—no live hostCanary4 sizes here
Canary 180m FlashNVIDIASpeech to textCanary 180m Flash is NVIDIA's tiny downloadable speech-to-text model built for raw speed rather than accuracy.93.7%2,484×40.2B—no live hostCanary 1b FlashNVIDIASpeech to textCanary 1b Flash is NVIDIA's tiny downloadable speech-to-text model that turns audio into written words.94.2%2,126×41B—no live hostCanary 1bNVIDIASpeech to textCanary 1b is NVIDIA's tiny downloadable speech-to-text model that turns audio into written words.94.2%766×41B—no live hostCanary 1b v2NVIDIASpeech to textCanary 1b v2 is a small downloadable speech-to-text model from NVIDIA that turns audio into written text.93.6%1,821×251B—no live hostResonant2 sizes here
Resonant 1Reson8Speech to textResonant 1 is a proprietary speech-to-text model from Reson8 that turns recorded audio into written words.95.4%—9——no live hostResonant 1 FlashReson8Speech to textResonant 1 Flash is Reson8's proprietary speech-to-text model that handles nine languages, though accuracy has only been measured for English.95.3%—9——no live hostMoonshine4 sizes here
Moonshine Streaming TinyUseful SensorsSpeech to textMoonshine Streaming Tiny is a 30-million-parameter English speech-to-text model that trades accuracy for extreme speed.88.8%4,375×130M—no live hostMoonshine Streaming MediumUseful SensorsSpeech to textMoonshine Streaming Medium is a tiny downloadable speech-to-text model from Useful Sensors that turns audio into written words.94.2%2,681×10.2B—no live hostMoonshine TinyUseful SensorsSpeech to textMoonshine Tiny is a 30-million-parameter speech-to-text model that turns audio into written words under a permissive MIT licence.88.6%3,733×130M—no live hostMoonshine Streaming SmallUseful SensorsSpeech to textMoonshine Streaming Small is a tiny English speech-to-text model from Useful Sensors, weighing roughly 0.1 billion parameters under a permissive MIT licence.93.1%3,206×10.1B—no live hostUniversal2 sizes here
Universal 3 ProAssemblyAISpeech to textUniversal 3 Pro is AssemblyAI's hosted-only speech-to-text model that converts audio to written words across 99 languages.94.8%—99——no live hostUniversal 3 5 ProAssemblyAISpeech to textUniversal 3 5 Pro is AssemblyAI's hosted-only speech-to-text model that turns audio into written words.95%—18——no live hostAudio8 ASR 0.1BAutoArk AISpeech to textAudio8 ASR is a tiny downloadable speech-to-text model from AutoArk AI that can transcribe an hour of audio in about five seconds.93%709×70.3B—no live hostZipformer cr CTC Transducer XL 290MSounds Good AISpeech to textZipformer cr CTC Transducer XL is a tiny downloadable speech-to-text model that turns English audio into written words at roughly 160 times real time.94.8%160×10.3B—no live hostMOSS Transcribe Preview 2BOpenMOSSSpeech to textMOSS Transcribe Preview is a downloadable speech-to-text model from OpenMOSS that turns audio into written words.95.3%149×12.4B—no live hostARK ASR2 sizes here
ARK ASR 3BAutoArk AISpeech to textARK ASR is a compact downloadable speech-to-text model from AutoArk AI that turns audio into written words.95.4%484×194.1B—no live hostARK ASR 0.6BAutoArk AISpeech to textARK ASR is a tiny downloadable speech-to-text model from AutoArk AI that turns audio into written words across 19 languages.94.9%672×191.3B—no live hostHojo ASR V1Hojo AISpeech to textHojo ASR V1 is a 5.2-billion-parameter speech-to-text model you can download and run yourself, licensed under Apache 2.0.95.5%73.8×25.2B—no live hostMOSS Transcribe DiarizeOpenMOSSSpeech to textMOSS Transcribe Diarize is a sub-1-billion-parameter speech-to-text model from OpenMOSS that identifies who is speaking while it transcribes.94.8%382×20.9B—no live hostZipformer Transducer XL 290MSounds Good AISpeech to textZipformer Transducer XL is a tiny downloadable speech-to-text model from Sounds Good AI that turns English audio into written words.93.8%141×10.3B—no live hostHiggs Audio2 sizes here
Higgs Audio v3 8b STT v2Boson AISpeech to textHiggs Audio is an English-only speech-to-text model from Boson AI that turns audio into written words at exceptional speed.95.3%139×18.9B—no live hostHiggs Audio v3 STTBoson AISpeech to textHiggs Audio v3 STT is a 2.68-billion-parameter English speech-to-text model from Boson AI with a permissive Apache licence.95.4%110×12.7B—no live host1 / 2