MODEL SIGNAL · MICROSOFT · NEW
MAI-Voice-2.1
An expressive, low-latency speech model for natural-sounding text-to-speech, supporting real-time and long-form generation.
CATEGORYMultimodal
RELEASEDOctober 1, 2026
Key Features
- Expressive text-to-speech generation
- Low-latency real-time and long-form speech generation
- Support for 23 languages and 26 locales with consistent voice identity
- Granular emotion control
- Zero-shot voice prompting
- Available through Microsoft Foundry and Azure Speech
Read the Model Signal report →
The designed Model Signal report for MAI-Voice-2.1 is still in review. It will publish here after the fact and quality gates clear.