← News
News

Microsoft open-sources MAI-Voice-2.1 and MAI-Transcribe-2-Streaming: 23/60 native accents, 150ms/100ms latency

Microsoft's MAI team released two open-source speech-stack models on October 3: MAI-Voice-2.1 (23 languages with native accents, Flash variant 150 ms first-packet latency) and MAI-Transcribe-2-Streaming (real-time transcription in 60 languages, ~100 ms first-batch latency).

Source: AI 日报(ai6s.net)

What it is

On October 3, Microsoft's MAI team released two open-source speech-stack models:

  • MAI-Voice-2.1: TTS with 23 languages of native accents (not generic-English-to-local), Flash variant at ~150 ms first-packet latency
  • MAI-Transcribe-2-Streaming: real-time transcription in 60 languages, ~100 ms first-batch latency, streaming output

Comparison

Against Whisper-large-v3 / CosyVoice and other mainstream open speech stacks:

Dimension MAI-Voice-2.1 CosyVoice-2 Whisper-large-v3
Realtime streaming ✓ (Flash 150ms) ✓ (partial) ✗ (offline mostly)
Native accents 23 Chinese dialects mostly N/A (recognition only)
First-packet latency 150ms ~200ms N/A
License open-source open-source (community) MIT
Dimension MAI-Transcribe-2-Streaming Whisper-large-v3 Paraformer
Streaming ✓ ✗ ✓
Languages 60 99 ZH/EN
First-batch latency 100ms N/A ~300ms

Local deployment angle

Flash-variant TTS targets realtime dialogue / customer-service scenarios — runs on a single consumer GPU. Streaming ASR is sensitive to memory and CPU scheduling — 16 GB+ workstation or cloud inference recommended. Watch community quantization and ONNX exports.