Models / Meta/ Muse Voice Transcribe

Muse Voice Transcribe

Meta

Speech to textTranscribes a recording into words

Input: audio. Output: text.InputOutput
Type
Closed
Input
None held
Output
None held
Cached
None held

We don't hold a list price for this model yet · hosted only — no weights published

Our take

Written Sep 14, 2026

Muse Voice Transcribe is Meta's proprietary speech-to-text model covering 25 languages. It excels at meeting-room and accented audio where most models struggle, yet falls behind on clean read-aloud and everyday internet video.

Who should pick it

Pick this for meeting transcription with overlapping speakers and room microphones, or content with accented English speakers. Use it in multilingual pipelines needing 25 languages where English accuracy is the priority. Skip it if you need clean read-aloud precision, everyday podcast or YouTube transcription, or any commercial access path — we list no live offers and no published licence.

The case for it

  • Among the better options for meeting-room audio with overlapping speech: 10.3% word error rate, better than most of 85 models.
  • Handles accented speech better than most: 5.8% word error rate on earnings calls, against a field middle of 7.9%.
  • 25 languages supported, broader than many transcription models.

The case against it

  • Weaker than most on clean read-aloud audio — 2.1% word error rate against a field middle of 1.5%.
  • Weaker than most on everyday internet audio like podcasts and YouTube — 8.9% word error rate against a field middle of 8.3%.
  • No commercial access path: proprietary weights, no published licence, and zero live offers in our data.
00

How good is it?

A speech-to-text model for transcribing recordings, with a particular edge on accented speakers.

Good at
  • transcribing speakers with a range of accentsAccented speech · 13th of 76
Less good at
  • transcribing clear recordings of people reading aloudClean read speech · 70th of 92

TranscriptionTurning speech into text3.5 of 5Open ASR WER · 28th of 76

Words it gets right

95.2%

Misses roughly one word in 21, averaged over nine English test sets.

Languages

25

Stated by the leaderboard; we do not hold the list itself.

Where it struggles
Read aloudaudiobooks, clean recording1.8%70th of 92
Podcasts and videoeveryday internet audio7.8%24th of 92
Accented speechspeakers from many countries6%13th of 76
Meetingsa room, several people, far microphone9.2%38th of 92

Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.

The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.

Other boards it appears on
European-accented speech 7th of 92Podcasts and video 24th of 92Recorded meetings 38th of 92Financial calls 65th of 92Clean read speech 70th of 92Harder read speech 70th of 92Accented speech 13th of 76

Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.

Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
4.84source ↗
6.01source ↗
3.21source ↗
9.2source ↗
2.02source ↗
7.75source ↗
1.84source ↗
4.51source ↗
01

Where to get it

We hold no priced listing for Muse Voice Transcribe.

There is no copy to download and no host in our price data, so Meta is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what Meta sells.

02

When we formed this view

Recent changes

Sep 21, 2026BenchmarkScored 4.84 on Open ASR WER
What movedleaderboard
Sep 21, 2026BenchmarkScored 6.01 on Accented speech
What movedleaderboard
Sep 21, 2026BenchmarkScored 3.21 on Financial calls
What movedleaderboard
Sep 21, 2026BenchmarkScored 9.2 on Recorded meetings
What movedleaderboard
Sep 21, 2026BenchmarkScored 2.02 on European-accented speech
What movedleaderboard
Sep 21, 2026BenchmarkScored 7.75 on Podcasts and video
What movedleaderboard
Sep 21, 2026BenchmarkScored 1.84 on Clean read speech
What movedleaderboard
Sep 21, 2026BenchmarkScored 4.51 on Harder read speech
What movedleaderboard
Sep 14, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • We don't hold a list price for this model yet — the gap is ours, not the lab's.
03

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence row of this model's own. A source states its weights are not published, so the determination that governs it is the one for closed weights, API access only.

Identifiers

Takes in, gives back
Audio in, text out
Catalogue slug
meta-muse-voice-transcribe

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us