Independent project. Not a U.S. government website.

USASI

Release

Microsoft AI releases MAI-Transcribe-2-Streaming and MAI-Voice-2.1 speech models

Event: · Published:

Microsoft AI announced three speech models on October 1, 2026. MAI-Transcribe-2-Streaming is a real-time transcription model that Microsoft says covers 60 languages with automatic, continuous language detection. MAI-Voice-2.1 is a multilingual text-to-speech model that Microsoft says supports 23 languages and 26 locales, and MAI-Voice-2.1-Flash is described as a faster variant with the same languages and features. Microsoft lists Microsoft Foundry, MAI Playground, and Vercel among the places all three can be used. The announcement does not mention published weights or a license.1

Sources

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project