Microsoft AI launched MAI-Transcribe-2-Streaming on October 1, its first model that transcribes speech live, as a person is still talking. It is in public preview in Microsoft Foundry at an introductory $0.54 per audio hour. Microsoft also released two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.
New to this? Read it in simple words
- Microsoft released an AI model that writes down speech while people are still talking.
- It works in 60 languages and is in a test version for developers.
- An outside test group found it made the fewest mistakes of 32 models shown.
- Microsoft also released two new models that turn text into spoken voice.
- Transcription
- Turning spoken words into written text.
- Streaming
- Working on speech bit by bit as it arrives, not after it ends.
- Public preview
- An early version that anyone can try, but that is not final.
What Microsoft launched
MAI-Transcribe-2-Streaming supports 60 languages and detects the language by itself, Microsoft says. Developers can use it in Microsoft Foundry, where its documentation says it is in public preview.
The model can power apps that show live captions, or start working on a request before the user has finished speaking, SiliconANGLE writes.
The two voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, speak 23 languages. They cost $22 and $15 per million characters.
Artificial Analysis’s streaming leaderboard, first-partial chart, as we read it on October 2.
What the outside test shows
Artificial Analysis tests streaming models on three sets of recordings. On its chart for the first partial transcript after speech ends, MAI-Transcribe-2-Streaming gets 2.5% of words wrong. The next model shown gets 3.4%.
Microsoft says the model also ranks first for final transcripts there. We checked the first-partial chart only.
The leaderboard shows 32 of 38 models in its default view.
Microsoft’s own claims
Microsoft says the model produces its first guesses “in just over 100ms”. In its internal tests, it says, words appear twice as fast as with its closest competitor, which it does not name.
SiliconANGLE and RuntimeWire also report a Microsoft figure of about 320 milliseconds for words to appear. Real response times also depend on the user’s network connection, they note.
The $0.54 per hour price is an introductory offer until the end of 2026, Microsoft and WinCentral say.
Sources
Every fact in this story comes from the sources below. Open them to check our work.
- 1Primary source · October 1, 2026Our first streaming transcription model debuts at no. 1 on Artificial Analysis Microsoft AI
- 2
- 3
- 4Research · October 1, 2026Microsoft targets ultra-realistic voice agents with its first streaming transcription model SiliconANGLE
- 5Research · October 1, 2026Microsoft adds streaming transcription and two voice models to its MAI lineup RuntimeWire
- 6Research · October 2, 2026Microsoft Launches New MAI Voice AI Models for Faster Real-Time Conversations WinCentral
We read Microsoft’s post and documentation, and compared SiliconANGLE, RuntimeWire and WinCentral. We opened Artificial Analysis’s streaming leaderboard ourselves on October 2 and read the first-partial error chart. Microsoft’s speed figures differ by measure (about 100 ms for first guesses, about 320 ms for words), and both are its own.