MAI-Voice-2.1

microsoft • mai-voice-2-1
Model Information
Slug mai-voice-2-1
LLMs.txt View
Release Date October 1, 2026 New
Organization
Model Description
MAI-Voice-2.1 is Microsoft AI's highest-fidelity, most expressive text-to-speech model. It produces natural, studio-grade speech across 23 languages, with detailed prosody, nuanced expressiveness, and speaker consistency over long-form content. It is suited for audiobooks, podcasts, lectures, narration, and brand audio where maximum voice quality matters. The model prioritizes naturalness and expressivity over latency-critical generation.

On OpenRouter, set `voice` to a full voice ID with the model suffix, such as `"en-US-Harper:MAI-Voice-2.1"`. A voice's locale sets the synthesis language. Set `response_format` to `"mp3"` or `"pcm"` (24 kHz mono). Harper and Grant support the `agent`, `customer-call-center`, `educational`, and `narrator` speaking styles, and many locale voices add emotion styles such as `excited`, `happy`, `sad`, and `whispering`. The full list of voices is in the `supported_voices` field of the [models API](https://openrouter.ai/api/v1/models?output_modalities=speech). See the [text-to-speech guide](https://openrouter.ai/docs/guides/overview/multimodal/tts).
Available at 2 Providers
Provider Type Model Name Original Model Input ($/1M) Output ($/1M) Free Actions
OpenRouter
OpenRouter
Chat Code
MAI-Voice-2.1
microsoft/mai-voice-2.1 $22.00 $0.00
Vercel AI Gateway
Vercel AI Gateway
MAI-Voice-2.1
microsoft/mai-voice-2.1 $22.00 -