Our first streaming transcription model debuts at no. 1 on Artificial Analysis | Microsoft AI
MAI-Transcribe-2-Streaming provides real-time transcription across 60 languages with 100ms partial hypothesis latency, priced at $0.54 per audio hour through year-end.
- MAI-Voice-2.1 supports 23 languages and 26 locales while preserving a single speaker's voice profile across languages for $22 per 1M characters.
- MAI-Voice-2.1-Flash generates 45 seconds of audio with 150ms end-to-end latency and costs $15 per 1M characters.
- Both voice models support voice cloning using a few seconds of reference audio with built-in consent guardrails.
- A live demonstration application called Chatter is available in the MAI Playground.
Developers building voice agents can achieve faster response loops and consistent cross-lingual voice profiles at lower latency and inference costs.

Sources
Read this as text
Back to the AI news