Modulate Vfast
Modulate
Speech to textTranscribes a recording into words
- Type
- Closed
- Input
- None held
- Output
- None held
- Cached
- None held
We don't hold a list price for this model yet · hosted only — no weights published
Our take
Written Sep 5, 2026Modulate Vfast is a hosted-only speech-to-text model that ranks among the most accurate we list on both clean read-aloud audio and meeting-room recordings. It also outperforms most rivals on podcasts, video and accented speech, though no commercial access path is currently tracked.
Choose this when accuracy on difficult audio matters most — it matches the best measured on meeting-room recordings and clean read-aloud, and beats most alternatives on podcasts, video and accented speech. Skip it if you need confirmed non-English support, self-hosting, or any pricing or availability information.
The case for it
- Matches the best measured on clean read-aloud audio at 0.9% word error, against a field median of 1.4%.
- Matches the best measured on meeting-room recordings at 6.7% word error, against a field median of 10.6%.
- Better than most on podcasts and video at 7.3% word error, well under the 8.2% median.
- Better than most on accented speech at 8.1% word error, against a 10.8% median.
The case against it
- No offers or pricing currently tracked — no commercial access path we can verify.
- Language coverage unverified; every accuracy figure is English only.
How good is it?
A speech-to-text model for meeting recordings, podcasts and clear read-aloud audio.
- transcribing recordings of meetings in a roomRecorded meetings · 3rd of 92
- transcribing podcasts and video audioPodcasts and video · 8th of 92
- transcribing clear recordings of people reading aloudClean read speech · 2nd of 92
EverydayGeneral questions and everyday reasoning
Not yet scored on Arena Text (overall).
CodingWriting and fixing code on its own
Not yet scored on Arena Coding.
AgenticPlanning, calling tools, staying on task
Not yet scored on Arena Agent.
WritingDrafting and rewriting prose
Not yet scored on Arena Creative Writing.
Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.
Every published score for this model6 scoresEvery figure we hold, from 6 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Where to get it
We hold no priced listing for Modulate Vfast.
There is no copy to download and no host in our price data, so Modulate is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what Modulate sells.
When we formed this view
Recent changes
What moved
first indexed by our pipelineEach date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- We don't hold a list price for this model yet — the gap is ours, not the lab's.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
We hold no licence row of this model's own. A source states its weights are not published, so the determination that governs it is the one for closed weights, API access only.
Identifiers
- Takes in, gives back
- Audio in, text out
- Catalogue slug
- modulate-vfast