Microsoft MAI-Voice-2.1 and MAI-Transcribe-2 Explained: Speech Models, Pricing and Use Cases

Microsoft MAI-Voice-2.1 and MAI-Transcribe-2 Explained: Speech Models, Pricing and Use Cases

MAI-Voice-2.1 and MAI-Transcribe-2: Microsoft’s Speech Models Explained

A voice agent on a plant-floor kiosk or in a contact centre is judged by one number the user feels but never sees: the gap between the moment they stop talking and the moment they hear a reply. On 1 October 2026 Microsoft AI shipped three models aimed squarely at that gap: MAI-Voice-2.1, a quality-focused text-to-speech model, MAI-Voice-2.1-Flash, a latency-optimised sibling, and MAI-Transcribe-2-Streaming, the company’s first streaming speech-to-text model. Together with the batch MAI-Transcribe-2 released in September, they complete a first-party speech stack that Microsoft sells through Foundry and Azure Speech.

The announcement is heavy on headline numbers and light on architecture. There is no model card with parameter counts, no training-data disclosure, and the weights are closed. What you can verify is the language coverage, the introductory pricing, the vendor-reported latency figures, the SSML and REST surface, and the preview status that rules out production commitments today.

What this covers: what Microsoft has confirmed and what it has not, how the three models fit a streaming voice pipeline, a worked latency budget and cost model (clearly labelled illustrative), a decision matrix against the alternatives, and a checklist for deciding whether to pilot now or wait for general availability.

Context and Background

Speech has been a Microsoft product line for years: Azure AI Speech has offered neural text-to-speech, speech-to-text and Custom Neural Voice under one resource. What changed in 2026 is that Microsoft AI, the consumer and frontier-model group, started shipping its own “MAI” family into Foundry alongside third-party models. The batch model MAI-Transcribe-2 appeared in September 2026 (the Foundry blog post is dated 3 September), and the 1 October release adds the streaming transcription variant and the 2.1 voice generation. Foundry is the delivery channel: you reach the models through a Microsoft Foundry resource for Speech, and Microsoft also lists availability through its own MAI Playground, Vercel, OpenRouter and, for the streaming transcriber, Azure Voice Live, with LiveKit shown as coming soon.

The competitive backdrop matters because speech is where agent products are won on feel rather than on benchmark scores. ElevenLabs pushed expressive synthesis hard; our own analysis of ElevenLabs Eleven v4 and its expressive TTS architecture covers the incumbent that every new entrant is measured against. On the recognition side, independent trackers such as Artificial Analysis now publish a streaming speech-to-text board that scores accuracy and settle time together, and that is the board on which Microsoft claims first place. Understanding what that board does and does not measure is half of reading this release correctly.

For an industrial or enterprise audience, the relevant question is less “which model sounds best” and more “can I put a bounded, budgetable, auditable voice interface in front of a system of record”. Digital-twin front ends, maintenance copilots and shop-floor assistants all share one constraint: the audio path must be fast enough to feel conversational, cheap enough to leave on, and governed enough to survive a security review. The MAI models are Microsoft’s answer on the first two axes, and the Azure surface (regions, keys, gated voice cloning) is the answer on the third. All three models carry a public-preview designation, which means no service-level agreement and, in Microsoft’s own wording on the Learn page for the voices, “not recommended for production workloads”.

I use the Microsoft AI announcement, the MAI-Voice-2.1 model page and the Azure Speech documentation for MAI voices as the primary sources throughout. Figures that appear only in secondary reporting are marked as such, and numbers Microsoft did not publish are called out as undisclosed rather than estimated.

The Reference Architecture: A Three-Model Voice Pipeline

A production voice agent built from these models is a three-stage streaming pipeline: MAI-Transcribe-2-Streaming turns incoming audio into partial and final transcripts over a WebSocket, a language model reasons over the text and calls tools, and MAI-Voice-2.1-Flash (or the standard 2.1 for pre-rendered content) synthesises the reply. Latency is dominated by how well you overlap the stages, not by any single model.

MAI-Voice-2.1 and MAI-Transcribe-2-Streaming voice agent pipeline reference architecture

Figure 1: Streaming voice agent pipeline. Caller audio flows through MAI-Transcribe-2-Streaming, a reasoning model with tools, and MAI-Voice-2.1-Flash back to the caller, with partial transcripts allowing the reasoning stage to start early.

The figure shows the data path and the one feedback edge that matters: barge-in. When the caller interrupts, the transcriber’s partials must cancel the in-flight synthesis, which means the TTS stage has to be cancellable mid-stream and the playback buffer on the client has to be flushable. None of that is a model feature; it is orchestration you write or buy, for example in a framework such as LiveKit or in Azure Voice Live.

MAI-Transcribe-2-Streaming: what is confirmed

Microsoft’s announcement states that the model supports 60 languages with automatic, continuous language detection, produces its first partial hypotheses “in just over 100ms of receiving audio”, and ranks first on Artificial Analysis for accuracy of both partial and final transcripts. It also claims that words appear in the transcript roughly twice as fast as the closest competitor in an internal evaluation, and that the model sits on the Pareto frontier of accuracy against latency.

Secondary coverage adds specifics that I treat as reported rather than confirmed from Microsoft’s own pages. These include a 2.5 percent word error rate for final transcripts at 0.13 seconds of settle time, the same 2.5 percent for first partials at 0.12 seconds, and a ranking of first among 38 models on the Artificial Analysis streaming index, measured on roughly eight hours of audio mixed from AA-AgentTalk (50 percent), VoxPopuli (25 percent) and Earnings22 (25 percent). One tracker commentary noted that the third-party board had not yet plotted the model at the time of writing, so the figure was effectively vendor-supplied. Another report quotes Microsoft’s materials as saying words appear “as early as 320 milliseconds in most cases”, which is a different metric from first-partial time and should not be conflated with it.

Delivery is described as incremental transcript updates over a WebSocket. Secondary reporting says the interface is compatible with the OpenAI Realtime API shape and that the Azure Speech SDK is also supported. I could not confirm the exact message schema from a Microsoft documentation page for the streaming model during this research, so treat the wire format as something to read from the Foundry documentation before you design against it.

Pricing is $0.54 per audio hour, an introductory rate that Microsoft says runs through the end of 2026. For comparison, the batch MAI-Transcribe-2 is reported at $0.10 per audio hour on a limited-time basis through the same date. The roughly five-fold premium for streaming is the cost of low settle time and incremental output, and it is a common pattern in the market.

MAI-Voice-2.1 and Flash: what is confirmed

Both voice models cover 23 languages across 26 locales. The model page lists English in four variants (US, Australia, UK, India), Italian, French, German, Hindi, Spanish (Spain, Mexico), Portuguese (Brazil, Portugal), Korean, Simplified Chinese, Turkish, Russian, Thai, Dutch, Romanian, Hungarian, Czech, Danish, Finnish, Indonesian, Polish, Swedish, Norwegian Bokmal and Vietnamese. The Learn documentation also mentions Japanese voices, so the exact language list differs slightly between pages; verify the locale you need in the voice table rather than trusting either summary.

Microsoft’s published model-inference latency is about 550 ms for MAI-Voice-2.1 to produce 45 seconds of audio and about 45 ms for MAI-Voice-2.1-Flash, with a stated end-to-end latency near 150 ms for Flash. The announcement describes Flash as 55 percent faster than 2.1 at inference and roughly 60 percent cheaper than comparable models, without naming the comparison set. List prices are $22 per million characters for 2.1 and $15 per million characters for Flash, also introductory through 31 December 2026 according to secondary reporting of Microsoft’s pricing note.

Both models support zero-shot voice prompting from a short reference clip, granular emotion control, and a single voice that keeps its identity across languages. The Learn page specifies instant voice cloning as a gated feature requiring Limited Access review, with a recommended reference duration of 5 to 60 seconds, and says only authorised, licensed voices are permitted in production. Microsoft positions 2.1 for audiobooks, content creation and voice-over, and Flash for call-centre agents, voice assistants and interactive voice response.

What Microsoft has not disclosed

Parameter counts, architecture family, training data, tokenizer or codec design, and context limits are not published for any of the three models. The Learn page does not state a maximum character count per request or a specific latency figure for Flash. Per-language accuracy for the 60 transcription languages is not broken out. Speaker diarization is documented for the batch model but, per the coverage I could find, not claimed for streaming. Anything beyond the list above is inference, and I flag it as such below.

Model family map for MAI-Voice-2.1, MAI-Voice-2.1-Flash and MAI-Transcribe-2 with the confirmed specs of each

Figure 2: The MAI speech family at a glance. Languages, price and latency as published or reported; undisclosed fields are deliberately absent rather than guessed.

The diagram is a map rather than an architecture: it records which numbers come from which model and where the boundary of public knowledge sits. That boundary is the honest summary of this release.

Deeper Analysis: Latency Budgets, Cost Models and the Azure Surface

The number that decides whether a voice agent feels natural is the turn gap: the time from the end of the user’s speech to the first audible byte of the reply. Humans leave roughly a few hundred milliseconds between turns in conversation, so every stage you add eats a visible share of that budget. The useful discipline is to write the budget down per stage and then ask which stages can overlap.

A worked latency budget (illustrative)

The table below combines Microsoft’s published or reported figures with assumptions I label explicitly. The model figures are vendor-reported; the language model, network and playback rows are placeholders for you to replace with measurements from your own deployment.

Stage Figure Source
End-of-speech to final transcript about 130 ms Reported, Microsoft claim via secondary coverage
Network, client to region and back 40 to 120 ms Assumption, depends on region and access network
LLM time to first token 250 to 600 ms Assumption, model and prompt dependent
TTS first audio, Flash end to end about 150 ms Microsoft stated
Client playback buffer start 20 to 60 ms Assumption
Total turn gap about 590 to 1,060 ms Illustrative sum

One secondary report summarises the full loop as roughly 580 to 780 milliseconds; that sits at the optimistic end of my range and implies a fast reasoning model close to the region. The point of the arithmetic is the shape, not the total: the transcription and synthesis stages together are under 300 ms of the budget, so the language model and the network are where you win or lose the remaining time. Choosing a smaller or cached reasoning path for routine intents often moves the needle more than swapping speech vendors.

Turn-gap latency budget sequence from end of speech to first audio with MAI-Transcribe-2-Streaming and MAI-Voice-2.1-Flash

Figure 3: Sequence of a single conversational turn. Speculative LLM start on stable partials overlaps recognition and reasoning, shortening the perceived gap.

Two overlap techniques matter. First, speculative generation: start the LLM on a stable partial transcript and discard the work if the final transcript differs materially. This trades extra token spend for lower latency, and it is only worthwhile when partials are accurate early, which is exactly the property Microsoft claims for this model (equal reported WER on first partials and finals). Second, sentence-level streaming into TTS: send the first clause to synthesis as soon as the LLM emits it, rather than waiting for the full response. Because Flash is priced per character, clause-level chunking does not change cost, but it does change prosody quality, since the synthesiser sees less context for each chunk.

Reading the 2.5 percent WER claim carefully

A word error rate is substitutions plus deletions plus insertions over reference words. A headline of 2.5 percent sounds like one error in forty words, but it is measured on a specific mix. The reported mix includes an agent-conversation set, parliamentary speech (VoxPopuli) and earnings calls (Earnings22), all relatively clean, close-microphone audio. Factory floors, vehicle cabs and speakerphones have lower signal-to-noise ratios, overlapping speakers and domain vocabulary such as part numbers, tag names and standards references. Expect the in-domain WER to be materially worse than the headline, and plan to measure it yourself.

A second subtlety is that streaming WER depends on when you score. A first partial at 0.12 seconds and a final at 0.13 seconds with the same WER is a striking claim, because in most streaming recognisers partials are revised as context arrives. If it holds on your data it simplifies the speculative-start pattern; if partials in your domain still revise heavily, you need a stability filter such as requiring the same prefix across two consecutive updates before acting. The independent comparison reported in the trade press put another leading model at 2.73 percent WER with a 0.49 second settle time, which illustrates the real finding: the claimed advantage is mostly latency, with a modest accuracy edge that sits within the noise of different test sets.

Cost model (illustrative)

Pricing is where the release is easiest to reason about because the units are simple: per audio hour for transcription, per million characters for synthesis. Take a hypothetical support line handling 10,000 four-minute calls a day. These are assumed figures for the arithmetic, not Microsoft data.

  • Streaming transcription of the full caller audio: 4 minutes is 0.0667 hours, times $0.54, is about $0.036 per call.
  • Agent speech of about two minutes at 150 words per minute is 300 words; at an assumed 6 characters per word including spaces that is 1,800 characters. Flash at $15 per million characters is about $0.027 per call; the standard 2.1 at $22 is about $0.040.
  • Speech total per call is therefore about $0.063 with Flash or $0.076 with 2.1, so roughly $630 or $760 per day at 10,000 calls, before the language model, telephony and orchestration costs.

Two observations follow. The speech layer at these rates is small next to the LLM tokens and telephony for most designs, so the 2.1 versus Flash choice is rarely about price; it is about latency and voice quality. And the introductory pricing expires at the end of 2026, so any budget you build now should include a sensitivity case where list prices rise. Microsoft’s Learn page points to the Azure Speech pricing page for the long-term numbers, and I could not confirm a post-introductory price.

The Azure surface: SSML, REST, regions and styles

The voices are consumed like other Azure Speech neural voices. The documented REST endpoint is a POST to https://{SPEECH_REGION}.tts.speech.microsoft.com/cognitiveservices/v1 with an application/ssml+xml body, an X-Microsoft-OutputFormat header such as audio-24khz-160kbitrate-mono-mp3, and a subscription key. SDKs are listed for Python, C#, JavaScript and Java. The voice name follows a locale-voice-model pattern; the documentation example is en-US-Harper:MAI-Voice-2.1-Flash, and the full voice table lists more than a hundred voice IDs across the 23 languages.

<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis"
       xmlns:mstts="http://www.w3.org/2001/mstts" xml:lang="en-US">
  <voice name="en-US-Harper:MAI-Voice-2.1-Flash">
    <mstts:express-as style="customer_call_center">
      Your work order has been scheduled for tomorrow morning.
    </mstts:express-as>
  </voice>
</speak>

Emotion and domain control uses mstts:express-as with a style attribute. The documentation lists styles such as angry, confused, determined, excited, happy, neutral, sad, shouting, softvoice and whispering, plus domain styles such as agent, audiobook, customer_call_center, educational and narrator, and states that availability varies by voice. That variation is a real engineering constraint: a style that exists on one voice may be rejected or ignored on another, so a voice-agent design should test its chosen voice against every style the dialogue manager might request.

Regions are limited. The Learn page states that 13 regions are supported but lists 14 region codes: France Central, East Asia, Southeast Asia, East US, Canada Central, East US 2, West US, West Europe, North Europe, West US 2, West US 3, Central India, Sweden Central and Japan East. The count discrepancy is a documentation inconsistency; verify against the portal. For data-residency conscious industrial customers, note the absence of some regions you may expect, and test that your Speech resource is created in a listed one.

Voice cloning is gated. Instant cloning needs a 5 to 60 second reference, no training step, and approval through the Limited Access review for Custom Neural Voice. Microsoft describes consent protections built into the system to prevent unauthorised cloning, and the documentation allows only authorised, licensed voices in production. For an enterprise this is a feature: a brand voice built on a consented speaker recording has a clean audit trail, and an unreviewed clone of a plant manager’s voice for internal announcements does not.

Decision flow for choosing MAI-Voice-2.1 versus MAI-Voice-2.1-Flash versus batch MAI-Transcribe-2 by latency and content type

Figure 4: Model selection flow. Interactive turns go to Flash and streaming transcription; long-form and offline content goes to MAI-Voice-2.1 and the cheaper batch transcriber.

Decision matrix

Use case Best fit in this release Alternative to evaluate Why
Shop-floor voice assistant, hands-busy MAI-Transcribe-2-Streaming plus Flash A noise-robust on-device recogniser Turn gap matters; in-domain noise unproven
Contact-centre agent in many locales Streaming plus Flash, 26 locales Incumbent TTS with broader language list Check your locale in the voice table first
Maintenance training narration, audiobooks MAI-Voice-2.1 Expressive-focused incumbents such as ElevenLabs Quality and speaker consistency over latency
Meeting and call analytics after the fact Batch MAI-Transcribe-2 Other batch ASR with diarization Lower price per hour; diarization documented for batch
Safety-critical command and control None of these today Deterministic push-to-talk, on-prem models Preview, no SLA, closed weights, no per-language WER

The last row is not a criticism of the models. It is a reminder that speech recognition errors, however rare, are the wrong failure mode to put in a path that actuates equipment. Voice belongs in the confirmation loop, never as the sole authority.

Trade-offs, Gotchas, and What Goes Wrong

Preview means preview. All three models are in public preview with no SLA. Microsoft’s own documentation says the feature is not recommended for production workloads and that some features may be unsupported or constrained. A pilot is reasonable; a customer-facing launch that depends on a specific latency percentile is not, because nothing contractual stands behind the numbers. Model versions can also change under a stable name during preview, which silently shifts both accuracy and cost.

Vendor-reported latency is a best case. The 130 ms, 150 ms and 45 ms figures describe model behaviour on Microsoft’s test setup. Your turn gap adds the network path, TLS and WebSocket setup, resampling, jitter buffers and telephony codecs. Telephony in particular narrows audio to 8 kHz, which changes recognition accuracy relative to the wideband clips used in public benchmarks. Measure at the 95th and 99th percentiles in the region you will deploy to, because tail latency is what users remember.

The benchmark mix is clean speech. An accuracy claim built from agent-conversation data, parliamentary recordings and earnings calls says little about reverberant halls, headset-less handheld devices or accented speech over a degraded link. Domain terms are the usual killer: asset tags, part numbers, chemical names and vendor product names. Without a custom vocabulary or phrase-boost feature (I could not confirm one for the streaming model), you may need a post-correction step that fuzzy-matches transcripts against a known entity list before the LLM sees them.

Language-count mismatches hide in the details. Sixty transcription languages and 23 synthesis languages are different sets. An agent that listens in Tamil but can only speak 23 languages, which are mostly European plus a few Asian ones, needs a fallback voice or a text channel. Microsoft’s pages also disagree slightly on whether Japanese is in the voice list, and on region counts, so confirm against the live voice table and portal.

Barge-in and echo are pipeline bugs, not model bugs. When the agent speaks, its own audio can leak into the microphone and be transcribed. Acoustic echo cancellation belongs on the client or in the telephony layer. When the user interrupts, you must cancel synthesis, flush audio already queued downstream, and truncate the conversation history to what the user actually heard, or the LLM will believe it said sentences the user never received.

Speculative decoding costs money. Starting the LLM on partials raises token spend because some starts are thrown away. If partial revision rates in your domain are 20 percent, you pay for about a quarter more prompt processing on speculative branches. That is usually a good trade for turn gap, but it should be a measured choice, not a default.

Closed weights and data handling. The models are closed, so there is no self-hosting, no quantisation, and no on-prem option. For sites with air-gapped networks or strict data-residency rules, audio leaves the premises to an Azure region. Confirm retention and logging terms in the Azure Speech data-privacy documentation before sending recordings that include customer or operational data. Voice cloning adds a further governance layer: a gated feature with a consent requirement is a safeguard, but it is also a process you must run.

Pricing is introductory. The $0.54 per hour, $22 and $15 per million character figures are stated as introductory through the end of 2026. Model the economics at a hypothetical higher price before committing, and remember that TTS cost scales with response verbosity: a chatty agent that doubles its average reply length doubles its synthesis bill, which is a prompt-engineering lever as much as a procurement one.

Practical Recommendations

Treat this release as an evaluation target for Q4 2026, not a production dependency. The strongest case is a latency-sensitive assistant where the turn gap is the product: Flash plus streaming transcription is a coherent pairing from a single vendor, with one billing relationship and one regional footprint. The weakest case is anything with a strict SLA, a safety function, or a need for languages outside the listed sets.

Run the evaluation in three passes. First, build a golden audio set from your own environment: at least a few hours of real recordings, with transcripts verified by a human, covering noise conditions, accents and your domain vocabulary. Score WER overall and on a domain-term subset, and compare the streaming model against your current recogniser on the same clips. Second, instrument the turn gap end to end with timestamps at each stage so you can see which hop dominates; do this from the actual client network, not a laptop on the data-centre LAN. Third, run listening tests on the voices you would deploy, with the dialogue styles your agent needs, and have native speakers rate naturalness for each locale.

Keep the architecture swappable. Put speech behind a thin interface (streaming recogniser in, text events out; text in, audio stream out) so that a preview model can be replaced if it regresses or is repriced. Log the model name and version string with every turn, because preview models change, and you will need that metadata when a regression appears.

  • Confirm your locales in the live voice table and portal, including Japanese or any other language the documentation treats inconsistently.
  • Create the Speech resource in a listed region and verify data-residency requirements.
  • Build a golden set of in-domain audio and measure WER, including on part numbers and tags, before trusting the 2.5 percent headline.
  • Measure p95 and p99 turn gap from the real client network, and set a latency budget per stage.
  • Implement barge-in with cancellable synthesis and echo cancellation before tuning anything else.
  • Model cost at introductory and at a higher hypothetical price, and cap agent reply length.
  • Apply for Limited Access review early if you need voice cloning, and document consent for every voice.
  • Keep voice out of any safety-critical actuation path; require confirmation for destructive commands.
  • Wrap the speech services behind an interface, and log model versions per turn.

Frequently Asked Questions

What is MAI-Voice-2.1?

MAI-Voice-2.1 is Microsoft AI’s quality-focused text-to-speech model, released on 1 October 2026 alongside a lower-latency variant, MAI-Voice-2.1-Flash. Both cover 23 languages across 26 locales, support granular emotion control through SSML, and offer zero-shot voice prompting from a short reference clip. Microsoft positions the standard model for audiobooks, voice-over and content creation. It is closed-weights, delivered through Microsoft Foundry and Azure Speech, in public preview with no service-level agreement.

How much do MAI-Voice-2.1 and MAI-Transcribe-2-Streaming cost?

The published introductory prices are $22 per million characters for MAI-Voice-2.1, $15 per million characters for MAI-Voice-2.1-Flash, and $0.54 per audio hour for MAI-Transcribe-2-Streaming, stated as running through the end of 2026. The batch MAI-Transcribe-2 is reported at $0.10 per audio hour. Microsoft’s documentation points to the Azure Speech pricing page for ongoing rates, and I could not confirm what prices apply after the introductory period ends, so budget with a margin.

How fast is MAI-Voice-2.1-Flash?

Microsoft publishes about 45 ms of model inference and about 150 ms end-to-end latency for MAI-Voice-2.1-Flash, and describes it as 55 percent faster than MAI-Voice-2.1, which needs about 550 ms to produce 45 seconds of audio. These are vendor figures from Microsoft’s own test setup. Your real turn gap also includes the network, the language model and playback buffering, so measure p95 and p99 latency from your deployment region before relying on them.

How many languages does MAI-Transcribe-2-Streaming support?

Microsoft states that MAI-Transcribe-2-Streaming supports 60 languages with automatic, continuous language detection, meaning it can follow a speaker who switches language mid-conversation. The synthesis models cover a smaller set of 23 languages and 26 locales. Microsoft has not published per-language accuracy for the 60 transcription languages, so the reported 2.5 percent word error rate should not be assumed to hold equally across all of them.

Is MAI-Transcribe-2-Streaming really the most accurate streaming speech-to-text model?

Microsoft says it ranks first on Artificial Analysis for both partial and final transcript accuracy. Secondary reports cite 2.5 percent WER at about 130 ms settle time, measured on roughly eight hours of clean benchmark audio. One analyst note observed that the independent board had not yet plotted the model, and that the lead over the next model is small on accuracy and large on latency. Treat it as a strong claim to validate on your own noisy, domain-specific audio.

Can I use MAI-Voice-2.1 in production or self-host it?

Not yet in either case. All three models are in public preview, which Microsoft describes as having no service-level agreement and not being recommended for production workloads. The weights are closed, so there is no self-hosting or on-premises option; you call the service in one of the supported Azure regions. Voice cloning is also gated behind a Limited Access review and requires consented, licensed voices. Plan pilots now and production only after general availability and contractual terms.

Further Reading

References

  1. Microsoft AI, “Our first streaming transcription model debuts at no. 1 on Artificial Analysis”, 1 October 2026. https://microsoft.ai/news/our-first-streaming-transcription-model/
  2. Microsoft AI, MAI-Voice-2.1 model page. https://microsoft.ai/models/mai-voice-2-1/
  3. Microsoft Learn, “MAI-Voice-2.1 and MAI-Voice-2.1-Flash” in Azure Speech documentation. https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices
  4. Microsoft Tech Community, “MAI-Transcribe-2: Highest quality transcription, at the fastest speed and lowest cost”, September 2026. https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/mai-transcribe-2-highest-quality-transcription-at-the-fastest-speed-and-lowest-c/4550972
  5. Microsoft Tech Community, “Build expressive voice experiences with new MAI models in Microsoft Foundry”, 1 October 2026. https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/build-expressive-voice-experiences-with-new-mai-models-in-microsoft-foundry/4524637
  6. MarkTechPost, coverage of MAI-Transcribe-2-Streaming, 2 October 2026 (secondary; source of the 2.5 percent WER and benchmark mix). https://www.marktechpost.com/2026/10/02/microsoft-ai-releases-mai-transcribe-2-streaming-1-real-time-speech-to-text-model-on-artificial-analysis/
  7. OrcaRouter, MAI-Transcribe-2-Streaming analysis (secondary; competitor latency comparison and undisclosed items). https://www.orcarouter.ai/blog/mai-transcribe-2-streaming-release

By Riju – about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *