← Back to Blog
02 October 2026 /// AI / VOICE-AGENTS / MICROSOFT

Microsoft's New Speech Model Writes the Caption While You Are Still Talking — and Debuts at No. 1

Microsoft's first streaming transcription model returns its first draft in just over 100 ms across 60 languages, tops Artificial Analysis for both partial and final accuracy, and costs $0.54 an audio hour through the end of 2026 — with two new voice models at $22 and $15 per million characters.

By Guruji Corporation
Infographic: Microsoft's streaming transcription model returns a first partial hypothesis in about 100 milliseconds, covers 60 transcription languages and 23 text-to-speech languages, and costs $0.54 per hour of audio
Original graphic by Guruji Corporation.

A voice bot that waits for you to finish your sentence is already having the wrong conversation. On 1 October 2026 Microsoft shipped three models built to end that wait, and its streaming speech-to-text model landed at the top of the third-party accuracy leaderboard on day one.

What actually happened

Microsoft announced its first streaming transcription model, plus two new voices. The company published MAI-Transcribe-2-Streaming alongside MAI-Voice-2.1 and MAI-Voice-2.1-Flash on 1 October 2026 — one speech-to-text model that listens continuously and two text-to-speech models that answer back, positioned explicitly as the building blocks for conversational voice agents.

The transcription model starts writing before you stop speaking. It streams across 60 languages with automatic, continuous language detection, and it produces its first hypotheses — "partial" transcripts — in just over 100 ms of receiving audio, revising them as more context arrives and committing a stable transcript when the utterance ends. Microsoft's own dictation and subtitling evaluations show words appearing in the transcript twice as fast as its closest competitor; that figure is Microsoft's internal evaluation, not an independent one.

It ranks first where ranking is independent. Microsoft claims the no. 1 position for accuracy on both final and partial transcripts on Artificial Analysis, and says it sits on the Pareto frontier of the accuracy-versus-latency chart. Artificial Analysis' streaming speech-to-text leaderboard does list MAI-Transcribe-2-Streaming at the top of its final-transcription ranking, measured on AA-WER Streaming — roughly eight hours of audio drawn from AA-AgentTalk, VoxPopuli and Earnings22, mixing accents, domain vocabulary and noisy conditions.

The two voice models split on purpose. MAI-Voice-2.1 supports 23 languages and 26 locales with a single voice that switches language mid-sentence and picks up a native accent rather than dragging one across — priced at $22 per million characters. MAI-Voice-2.1-Flash covers the same languages for high-volume, latency-sensitive work: 45 seconds of audio generated with an end-to-end latency of 150 ms, 55% faster inference and around 60% cheaper than what Microsoft calls comparable models, at $15 per million characters. Both support voice cloning from a few seconds of reference audio, with consent guardrails built in.

The price of listening is now a published hourly rate. MAI-Transcribe-2-Streaming carries an introductory price of $0.54 per hour of audio through the end of 2026 — about $9 per 1,000 minutes of recorded speech.

It is live, but it is a preview. Microsoft's own documentation puts the model in public preview: no service-level agreement, not recommended for production workloads, sessions capped at one hour per connection, audio sent over a WebSocket using an OpenAI-Realtime-style protocol, served from Sweden Central, Central US and — notably for Indian deployments — South India, with East US 2 marked "coming soon".

Why it matters for Indian SMBs and mid-market teams

  • Real-time voice just got a line item, not a project. At $0.54 an hour, live captioning or agent-assist on 10,000 minutes of support calls a month is roughly $90 — about ₹8,000 at current rates. That is cheap enough to put live transcription on ordinary phone calls instead of treating it as a premium feature.
  • Check your languages against the 23, not the 60. Transcription covers 60 languages; speech synthesis covers 23 languages across 26 locales. If your customers speak Hindi, Marathi, Tamil or Bengali and you want the agent to reply out loud, the synthesis list — not the transcription list — is the one that decides whether this works.
  • Preview status is an architecture constraint, not a footnote. No SLA and "not recommended for production" means you do not put this model in the only path a customer request can take. Keep a batch or fallback model behind it, and design for the one-hour session ceiling with clean reconnects.
  • Treat partials as provisional. A partial transcript is a working hypothesis; the model commits a final result separately. Anything irreversible — a payment, a ticket closure, a refund — should fire on the committed transcript, never on the partial that arrived mid-sentence.
  • Day-one benchmark leadership still needs your own audio. The ranking is on a public leaderboard built from mostly Western-accent datasets. Indian-English accents and Hindi-English code-switching are exactly where word error rates drift, so run your own recordings through it before you quote the number to anyone.

Where Guruji Corporation fits in

We build the unglamorous half of voice AI: the pipeline that takes a stream, decides what is final, routes it to the right system, logs the cost per call and keeps a fallback ready when the primary model is degraded or rate-limited. New model launches change which box in that diagram is cheapest or fastest — they rarely change the diagram itself, which is why a vendor swap should be an afternoon's benchmark rather than a quarter's migration.

If you are already paying for transcription, summarisation or an IVR, this week's useful hour is a measurement one: run a sample of your own call audio through the new model and compare error rate and cost against what you have. Talk to us and we will do that comparison honestly, tell you whether the switch is worth it, and wire the usage dashboards so the saving shows up as a number you can check at month end.


Sources

Want to see how we can build your project?

Send us your project outline or chat directly on WhatsApp.