Models / Microsoft/ Azure Speech 06 2026

Azure Speech 06 2026

Microsoft

Speech to textTranscribes a recording into words

Input: audio. Output: text.InputOutput
Type
Closed
Input
None held
Output
None held
Cached
None held

We don't hold a list price for this model yet · hosted only — no weights published

Our take

Written Sep 17, 2026

Microsoft's Azure Speech 06 2026 is a speech-to-text model with solid measured accuracy on English audio, better than most models on every recording condition we hold. The catch is that we list no download and no host for it, so there is no route here to actually run it.

Who should pick it

Reach for it when you are transcribing clearly recorded English audio — read-aloud material, podcasts, video — or English meeting recordings, where its error rate is better than most models measured on those conditions. Skip it if you need a route to actually run it, since we list no download and no host, or if you need a language other than English, speaker labelling, timestamps or streaming.

The case for it

  • On clean read-aloud recordings it gets 1.2% of words wrong, better than most models on that condition, so audiobooks and narrated material are where it looks strongest.
  • Podcast and video audio comes out at 7.3% of words wrong, better than most of the field there, which makes everyday internet audio a reasonable fit.
  • Meeting recordings land at 7% of words wrong, better than most models on that condition, though the whole field scores far worse on meetings than on clean speech.

The case against it

  • We list no download and no host for it, so there is nothing here you can try — the accuracy figures describe a model you cannot currently reach through this page.
  • Averaged over nine English test sets it sits at 4.3% of words wrong, above the middle of the field rather than at the top, where the best is 3.6%.
  • Every accuracy figure is English only: the model lists 25 languages, but none of them are measured here, so accuracy elsewhere is unverified.
00

How good is it?

A speech-to-text model for turning recordings of meetings, podcasts and varied accents into written text.

Good at
  • turning spoken English into written textOpen ASR WER · 8th of 76
  • transcribing recordings of meetings in a roomRecorded meetings · 5th of 92
  • transcribing speakers with a range of accentsAccented speech · 7th of 76
  • transcribing podcasts and video audioPodcasts and video · 9th of 92

TranscriptionTurning speech into text4 of 5Open ASR WER · 8th of 76

Words it gets right

95.7%

Misses roughly one word in 23, averaged over nine English test sets.

Languages

25

Stated by the leaderboard; we do not hold the list itself.

Where it struggles
Read aloudaudiobooks, clean recording1.2%24th of 92
Podcasts and videoeveryday internet audio7.3%9th of 92
Accented speechspeakers from many countries5.6%7th of 76
Meetingsa room, several people, far microphone7%5th of 92

Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.

The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.

Other boards it appears on
Recorded meetings 5th of 92Podcasts and video 9th of 92Harder read speech 18th of 92Clean read speech 24th of 92Financial calls 34th of 92European-accented speech 49th of 92Accented speech 7th of 76

Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.

Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
4.26source ↗
5.59source ↗
2.31source ↗
6.95source ↗
3.87source ↗
7.26source ↗
1.19source ↗
2.5source ↗
01

Where to get it

We hold no priced listing for Azure Speech 06 2026.

There is no copy to download and no host in our price data, so Microsoft is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what Microsoft sells.

02

Models people weigh against Azure Speech 06 2026

03

When we formed this view

Recent changes

Sep 11, 2026BenchmarkScored 4.26 on Open ASR WER
What movedleaderboard
Sep 11, 2026BenchmarkScored 5.59 on Accented speech
What movedleaderboard
Sep 11, 2026BenchmarkScored 2.5 on Harder read speech
What movedleaderboard
Aug 2, 2026BenchmarkScored 2.31 on Financial calls
What movedleaderboard
Aug 2, 2026BenchmarkScored 6.95 on Recorded meetings
What movedleaderboard
Aug 2, 2026BenchmarkScored 3.87 on European-accented speech
What movedleaderboard
Aug 2, 2026BenchmarkScored 7.26 on Podcasts and video
What movedleaderboard
Aug 2, 2026BenchmarkScored 1.19 on Clean read speech
What movedleaderboard
Aug 1, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • We don't hold a list price for this model yet — the gap is ours, not the lab's.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence row of this model's own. A source states its weights are not published, so the determination that governs it is the one for closed weights, API access only.

Identifiers

Takes in, gives back
Audio in, text out
Catalogue slug
microsoft-azure-speech-06-2026

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us