Microsoft's first streaming transcription model returns its first draft in just over 100 ms across 60 languages, tops Artificial Analysis for both partial and final accuracy, and costs $0.54 an audio hour through the end of 2026 — with two new voice models at $22 and $15 per million characters.
A voice bot that waits for you to finish your sentence is already having the wrong conversation. On 1 October 2026 Microsoft shipped three models built to end that wait, and its streaming speech-to-text model landed at the top of the third-party accuracy leaderboard on day one.
Microsoft announced its first streaming transcription model, plus two new voices. The company published MAI-Transcribe-2-Streaming alongside MAI-Voice-2.1 and MAI-Voice-2.1-Flash on 1 October 2026 — one speech-to-text model that listens continuously and two text-to-speech models that answer back, positioned explicitly as the building blocks for conversational voice agents.
The transcription model starts writing before you stop speaking. It streams across 60 languages with automatic, continuous language detection, and it produces its first hypotheses — "partial" transcripts — in just over 100 ms of receiving audio, revising them as more context arrives and committing a stable transcript when the utterance ends. Microsoft's own dictation and subtitling evaluations show words appearing in the transcript twice as fast as its closest competitor; that figure is Microsoft's internal evaluation, not an independent one.
It ranks first where ranking is independent. Microsoft claims the no. 1 position for accuracy on both final and partial transcripts on Artificial Analysis, and says it sits on the Pareto frontier of the accuracy-versus-latency chart. Artificial Analysis' streaming speech-to-text leaderboard does list MAI-Transcribe-2-Streaming at the top of its final-transcription ranking, measured on AA-WER Streaming — roughly eight hours of audio drawn from AA-AgentTalk, VoxPopuli and Earnings22, mixing accents, domain vocabulary and noisy conditions.
The two voice models split on purpose. MAI-Voice-2.1 supports 23 languages and 26 locales with a single voice that switches language mid-sentence and picks up a native accent rather than dragging one across — priced at $22 per million characters. MAI-Voice-2.1-Flash covers the same languages for high-volume, latency-sensitive work: 45 seconds of audio generated with an end-to-end latency of 150 ms, 55% faster inference and around 60% cheaper than what Microsoft calls comparable models, at $15 per million characters. Both support voice cloning from a few seconds of reference audio, with consent guardrails built in.
The price of listening is now a published hourly rate. MAI-Transcribe-2-Streaming carries an introductory price of $0.54 per hour of audio through the end of 2026 — about $9 per 1,000 minutes of recorded speech.
It is live, but it is a preview. Microsoft's own documentation puts the model in public preview: no service-level agreement, not recommended for production workloads, sessions capped at one hour per connection, audio sent over a WebSocket using an OpenAI-Realtime-style protocol, served from Sweden Central, Central US and — notably for Indian deployments — South India, with East US 2 marked "coming soon".
We build the unglamorous half of voice AI: the pipeline that takes a stream, decides what is final, routes it to the right system, logs the cost per call and keeps a fallback ready when the primary model is degraded or rate-limited. New model launches change which box in that diagram is cheapest or fastest — they rarely change the diagram itself, which is why a vendor swap should be an afternoon's benchmark rather than a quarter's migration.
If you are already paying for transcription, summarisation or an IVR, this week's useful hour is a measurement one: run a sample of your own call audio through the new model and compare error rate and cost against what you have. Talk to us and we will do that comparison honestly, tell you whether the switch is worth it, and wire the usage dashboards so the saving shows up as a number you can check at month end.
Sources
Send us your project outline or chat directly on WhatsApp.