0
MODEL SIGNAL · MICROSOFT · NEW

MAI-Voice-2.1

An expressive, low-latency speech model for natural-sounding text-to-speech, supporting real-time and long-form generation.

CATEGORYMultimodal
RELEASEDOctober 1, 2026
Key Features
  • Expressive text-to-speech generation
  • Low-latency real-time and long-form speech generation
  • Support for 23 languages and 26 locales with consistent voice identity
  • Granular emotion control
  • Zero-shot voice prompting
  • Available through Microsoft Foundry and Azure Speech

Provider announcement →

Read the Model Signal report →

The designed Model Signal report for MAI-Voice-2.1 is still in review. It will publish here after the fact and quality gates clear.