Models / AssemblyAI/ Universal 3 5 Pro

Universal 3 5 Pro

AssemblyAI

Speech to textTranscribes a recording into words

Input: audio. Output: text.InputOutput
Type
Proprietary
Input
None held
Output
None held
Cached
None held

We don't hold a list price for this model yet · hosted only — no weights published

Our take

Written Aug 2, 2026

Universal 3 5 Pro is AssemblyAI's hosted-only speech-to-text model that turns audio into written words. It achieves very low error rates on clean, structured English audio but degrades sharply in meetings or with accented speech, and no pricing or hosting options are currently tracked.

Who should pick it

Pick this for high-accuracy transcription of clean read-aloud English, financial calls, or European-accented English speech where the error rate stays under 2%. Skip it if you need to transcribe meetings, heavily accented speech, or podcasts and video, where the error rate rises nearly tenfold; also skip it if you need confirmed pricing, hosted access, or measured speed data.

The case for it

  • Excellent accuracy on clean, structured audio: about one word in ninety wrong on read-aloud English, and similarly strong on financial calls and European-accented speech.
  • 18 languages listed, though every accuracy figure we hold is English-only.

The case against it

  • Accuracy collapses in natural, unstructured conditions: recorded meetings see nearly ten times the error rate of clean read speech, and podcasts and video see nearly seven times.
  • No pricing or hosted access is currently tracked in our data.
  • Speed is explicitly unmeasured, so no latency claims can be made.
00

How good is it?

TranscriptionTurning speech into text4 of 5Open ASR WER · 16th of 74

Words it gets right

95%

Misses roughly one word in 20, averaged over nine English test sets.

Languages

18

Stated by the leaderboard; we do not hold the list itself.

Where it struggles
Read aloudaudiobooks, clean recording1.1%
Podcasts and videoeveryday internet audio7.6%
Accented speechspeakers from many countries10%
Meetingsa room, several people, far microphone10.6%

Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case, not to the board.

The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.

Also scored, on boards we give no mark for
European-accented speech 4th of 74Harder read speech 6th of 74Financial calls 9th of 74Clean read speech 12th of 74Podcasts and video 15th of 74Accented speech 24th of 74Recorded meetings 38th of 74

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.

Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
5independentsource ↗
10independentsource ↗
1.8independentsource ↗
10.6independentsource ↗
7.6independentsource ↗
1.1independentsource ↗
2.1independentsource ↗
01

Where to get it

We hold no priced listing for Universal 3 5 Pro.

There is no copy to download and no host in our price data, so AssemblyAI is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what AssemblyAI sells.

02

Models people weigh against Universal 3 5 Pro

03

When we formed this view

Dates behind this page

Aug 2, 2026BenchmarkScored 10 on Accented speechleaderboard
Aug 2, 2026BenchmarkScored 1.8 on Financial callsleaderboard
Aug 2, 2026BenchmarkScored 10.6 on Recorded meetingsleaderboard
Aug 2, 2026BenchmarkScored 1.9 on European-accented speechleaderboard
Aug 2, 2026BenchmarkScored 7.6 on Podcasts and videoleaderboard
Aug 2, 2026BenchmarkScored 1.1 on Clean read speechleaderboard
Aug 2, 2026BenchmarkScored 2.1 on Harder read speechleaderboard
Aug 1, 2026ListedListed on LLMapfirst indexed by our pipeline
Aug 1, 2026BenchmarkScored 5 on Open ASR WERleaderboard

Prices last checked 6h ago

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • We don't hold a list price for this model yet — the gap is ours, not the lab's.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

Commercial API terms. We hold no licence record for this model, so there is nothing to summarise here.

Identifiers

Modality record
audio->text
Catalogue slug
assemblyai-universal-3-5-pro

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us