Models / Zoom/ Scribe v1

Scribe v1

Zoom

Speech to textTranscribes a recording into words

Input: audio. Output: text.InputOutput
Type
Proprietary
Input
None held
Output
None held
Cached
None held

We don't hold a list price for this model yet · hosted only — no weights published

Our take

Written Aug 2, 2026

Scribe v1 is Zoom's proprietary speech-to-text model, available only through ElevenLabs, that turns English audio into written text. It is accurate on clean read-aloud recordings but its error rate rises sharply on accented speech, meetings and podcasts.

Who should pick it

Pick this for high-quality transcription of clean, scripted English audio where the recording is good. Use it for budget-conscious English speech-to-text work via ElevenLabs. Skip it if your audio is accented, multi-speaker, or from meetings or podcasts, or if you need any language other than English.

The case for it

  • Strong on clean read-aloud English, with about one word in ninety wrong, and similarly low error rates on financial calls and harder read speech.
  • Input cost at ElevenLabs is extremely low for the speech-to-text category.

The case against it

  • Accuracy collapses on challenging real-world audio: the error rate jumps more than eightfold from clean read-aloud to accented speech, and more than sixfold to podcasts, video and meetings.
  • Only one language is supported, and every accuracy figure we hold is English-only.
  • A single provider offer with no open weights or alternative hosts, so there is no portability if terms or pricing change.
00

How good is it?

TranscriptionTurning speech into text4.5 of 5Open ASR WER · 8th of 74

Words it gets right

95.3%

Misses roughly one word in 21, averaged over nine English test sets.

Languages

1

Stated by the leaderboard; we do not hold the list itself.

Where it struggles
Read aloudaudiobooks, clean recording1.1%
Podcasts and videoeveryday internet audio7.9%
Accented speechspeakers from many countries9.2%
Meetingsa room, several people, far microphone6.9%

Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case, not to the board.

The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.

Also scored, on boards we give no mark for
Financial calls 1st of 74Recorded meetings 2nd of 74Harder read speech 4th of 74Clean read speech 16th of 74Accented speech 18th of 74Podcasts and video 23rd of 74European-accented speech 42nd of 74

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.

Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
4.7independentsource ↗
9.2independentsource ↗
1.4independentsource ↗
6.9independentsource ↗
7.9independentsource ↗
1.1independentsource ↗
2.1independentsource ↗
01

Or rent it from someone else

Cheapest published offer

Cheapest of 1 live listings. Picked at the widest standard context we hold, within one quantisation slice, so the numbers beside it are a price one host actually charges.

per minute of audio
$0.004
Context served
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderPrice per minute of audioContextThroughputTrains on promptsLogs promptsZero retention
ElevenLabs$0.004not reportednot measuredUnknownUnknownUnknown

Across the 1 listings we hold: 0 say they do not train on prompts, 0 say they do and 1 do not say. 0 appear in the zero-retention registry we check; the rest are unknown to us rather than confirmed either way.

02

When we formed this view

Dates behind this page

Aug 2, 2026BenchmarkScored 9.2 on Accented speechleaderboard
Aug 2, 2026BenchmarkScored 1.4 on Financial callsleaderboard
Aug 2, 2026BenchmarkScored 6.9 on Recorded meetingsleaderboard
Aug 2, 2026BenchmarkScored 4 on European-accented speechleaderboard
Aug 2, 2026BenchmarkScored 7.9 on Podcasts and videoleaderboard
Aug 2, 2026BenchmarkScored 1.1 on Clean read speechleaderboard
Aug 2, 2026BenchmarkScored 2.1 on Harder read speechleaderboard
Aug 1, 2026ListedListed on LLMapfirst indexed by our pipeline
Aug 1, 2026BenchmarkScored 4.7 on Open ASR WERleaderboard

Prices last checked 6h ago

What we do not know about this model yet

  • 1 of 1 listings publish no parameter list, so what their API accepts is unknown to us.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 1 listings do not say whether they train on prompts.
  • We don't hold a list price for this model yet — the gap is ours, not the lab's.
  • We hold no cached-input rate for any of its listings.
03

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

Commercial API terms. We hold no licence record for this model, so there is nothing to summarise here.

Identifiers

Modality record
audio->text
Catalogue slug
zoom-scribe-v1

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us