MAI-Transcribe-2-Streaming: 100ms voice transcripts, $0.54/hr
Microsoft's MAI-Transcribe-2-Streaming returns partial transcripts in about 100ms for $0.54 an hour. Real-time speech-to-text costs 5x batch. Route accordingly.
Microsoft AI shipped MAI-Transcribe-2-Streaming on October 1, its first real-time speech-to-text model. It starts returning partial transcripts just over 100 milliseconds after it hears audio, covers 60 languages, and costs $0.54 per hour of audio. That is the number to read twice: the batch version, MAI-Transcribe-2, costs $0.10. Real-time transcription is now fast enough for a phone agent, and it costs more than five times as much as transcribing the same call afterward.
What actually happened
Per Microsoft AI's announcement and heise's coverage, Microsoft released three audio models at once:
- MAI-Transcribe-2-Streaming: speech-to-text with partial hypotheses about 100ms in, refined as more audio arrives. 60 languages with continuous language detection, so a caller who switches languages mid-call does not break it. Microsoft says it ranks first on Artificial Analysis for both partial and final transcripts.
- MAI-Voice-2.1: text-to-speech in 23 languages at $22 per million characters, with voice cloning from short reference audio. Heise reports cloning needs a Microsoft review and an uploaded consent recording.
- MAI-Voice-2.1-Flash: the low-latency voice for live agents, at $15 per million characters.
The $0.54 rate is introductory and runs through the end of 2026. Microsoft has not published what it costs after that. The models are live in Microsoft Foundry, the MAI Playground, Vercel, OpenRouter and Azure Voice Live, with LiveKit support listed as coming soon.
Why streaming transcription matters for your business
Latency is what makes a voice agent feel human. A phone agent has to hear, think and answer inside the pause a caller expects. Partial transcripts at 100ms let the agent start a lookup or a tool call before the caller finishes the sentence. That is the gap between "sounds like a bot" and "sounds like the front desk."
Do not stream what can wait. Live calls need streaming. Voicemail, call summaries, QA scoring and CRM notes do not. Send those through the batch model at $0.10 an hour. A shop with 500 hours of calls a month pays about $270 to stream all of it and about $50 to batch it. Split the pipeline by job, not by vendor.
Budget for the price after December. Both MAI transcription rates are labeled limited-time with no posted successor. Put the speech provider behind one adapter in your stack, normalize the transcript format, and keep a second vendor tested. If January brings a new rate card, you change a config line, not a pipeline.
Get consent on cloned voices in writing. If you clone a staff member's voice for your agent, keep the consent recording and the signed release. Microsoft asks for it, and so will your lawyer.
Key takeaways
- MAI-Transcribe-2-Streaming launched October 1, 2026 with partial transcripts in about 100ms
- $0.54 per audio hour through the end of 2026, vs $0.10 for the batch model
- 60 languages with continuous language detection; live on Foundry, Vercel and OpenRouter
- Stream live calls only; batch everything that can wait
- Both rates are introductory, so keep the provider swappable
Want a voice agent on your phone line? We build them with streaming where the caller is live, batch where they are not, and a swappable speech vendor. See what we've shipped or tell us about your call volume.
Sources: Microsoft AI, heise online.
- #speech-to-text
- #voice-agents
- #microsoft
- #real-time-transcription
- #ai-pricing
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Inworld buys Ultravox: your voice agent vendor just changed
Inworld acquired Ultravox, the real-time voice agent platform. Inworld voices got a free TTS-2 upgrade; third-party voices got 'nothing changes today.'
Read itGoogle WikiSkill: an agent failure wiki lifts small models
Google Research's WikiSkill gives AI agents a persistent wiki of past failures and fixes. Smaller models with it rivaled bigger models without it. Copy the pattern.
Read it