Introduction

On October 1, 2026, Microsoft AI shipped a same-day voice-agent trio: MAI-Transcribe-2-Streaming (its first streaming ASR), plus two closed TTS models — MAI-Voice-2.1 and the low-latency twin MAI-Voice-2.1-Flash. Primary source: Microsoft AI — Our first streaming transcription model debuts at no. 1 on Artificial Analysis (published Oct 1, 2026). Model cards: MAI-Voice-2.1.

Closed / API only — no open weights. You can run them today through Microsoft Foundry, the MAI Playground (including the Chatter demo), OpenRouter (both voice models), Vercel, and Azure Voice Live (LiveKit listed as coming soon).

Pro Tip

Treat the stack as a latency budget, not three isolated toys. Pair MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash when the agent must hear and answer inside a conversational window; keep MAI-Voice-2.1 for audiobook / voice-over takes where ~550 ms model inference and higher fidelity beat Flash’s ~45 ms inference / ~150 ms end-to-end claim.

What Shipped

Concrete facts from Microsoft’s Oct 1 post and model page:

ModelRoleConcrete fact

MAI-Transcribe-2-Streaming

Streaming ASR

#1 final + partial accuracy on Artificial Analysis streaming board; first partials just over ~100 ms; 60 languages with continuous auto language detection; intro price $0.54 / audio hour through end of 2026

MAI-Voice-2.1

Expressive TTS

23 languages / 26 locales; one voice can switch languages with a native accent; $22 / 1M characters; ~550 ms model inference

MAI-Voice-2.1-Flash

Agent TTS

Same language/locale set; up to 45 s audio; end-to-end latency about 150 ms; Microsoft claims ~55% faster inference and ~60% cheaper vs comparable models; $15 / 1M characters

Both voice models support zero-shot cloning from a few seconds of reference audio with built-in consent guardrails. Microsoft also cites a 4,000-listener Turing-style test where 50.3% of listeners rated MAI-Voice as equally or more human-like than human recordings — treat that as vendor-reported, then A/B on your own scripts.

This is separate from earlier MAI-Transcribe-2 batch ASR (September, billed around $0.10 / hour as a limited offer) and from prior MAI-Voice-2 / Voice-2-Flash cards on Microsoft Learn — 2.1 is the Oct 1 multilingual / pricing refresh.

Why Voice Makers Care

A voice agent is a loop: hear → decide → speak. Streaming partials let the app start tool calls or captions before the speaker finishes; Flash buys the reply side back. Multilingual “one speaker, many locales” is the other unlock — tutoring and support flows can keep the same persona across English, Mandarin, German, and the rest of the 23-language set without casting a new actor per market.

OpenRouter already lists both TTS models (dated October 1, 2026) at the same $22 / $15 per-million-character prices, so you are not locked into a single gateway.

How to Try It

  1. Read the Oct 1 Microsoft AI launch post and open the MAI-Voice-2.1 model page
  2. In the MAI Playground, try Chatter for a live agent loop, or hit the models via Microsoft Foundry
  3. For a quick TTS smoke test, call MAI-Voice-2.1 / MAI-Voice-2.1-Flash through OpenRouter or Vercel’s AI Gateway
  4. For streaming ASR, wire MAI-Transcribe-2-Streaming over the WebSocket path (Azure-backed listings show $0.54 / hour) and log partial vs final timestamps on a real call recording

# Direction sketch for MAI-Voice (emotion / style per Microsoft docs + playground)[warm, clear] Thanks for holding — I've got your order open.[soft laugh] That double charge on the eleventh? Refund already started.

Conclusion

Oct 1’s MAI drop is the clearest Microsoft audio product launch in this window: a chart-topping streaming ASR at $0.54/hr, plus a fidelity/latency TTS pair at $22 / $15 per 1M characters, runnable the same day on Foundry, OpenRouter, and the MAI Playground. Start from the official post, A/B Flash vs 2.1 on a familiar script, and keep Eleven v4 / Cartesia Sonic in their own lanes when you compare expressiveness.

—Anabel ♡