News
Microsoft open-sources MAI-Voice-2.1 and MAI-Transcribe-2-Streaming: 23/60 native accents, 150ms/100ms latency
Microsoft's MAI team released two open-source speech-stack models on October 3: MAI-Voice-2.1 (23 languages with native accents, Flash variant 150 ms first-packet latency) and MAI-Transcribe-2-Streaming (real-time transcription in 60 languages, ~100 ms first-batch latency).
Source: AI 日报(ai6s.net)
What it is
On October 3, Microsoft's MAI team released two open-source speech-stack models:
- MAI-Voice-2.1: TTS with 23 languages of native accents (not generic-English-to-local), Flash variant at ~150 ms first-packet latency
- MAI-Transcribe-2-Streaming: real-time transcription in 60 languages, ~100 ms first-batch latency, streaming output
Comparison
Against Whisper-large-v3 / CosyVoice and other mainstream open speech stacks:
| Dimension | MAI-Voice-2.1 | CosyVoice-2 | Whisper-large-v3 |
|---|---|---|---|
| Realtime streaming | ✓ (Flash 150ms) | ✓ (partial) | ✗ (offline mostly) |
| Native accents | 23 | Chinese dialects mostly | N/A (recognition only) |
| First-packet latency | 150ms | ~200ms | N/A |
| License | open-source | open-source (community) | MIT |
| Dimension | MAI-Transcribe-2-Streaming | Whisper-large-v3 | Paraformer |
|---|---|---|---|
| Streaming | ✓ | ✗ | ✓ |
| Languages | 60 | 99 | ZH/EN |
| First-batch latency | 100ms | N/A | ~300ms |
Local deployment angle
Flash-variant TTS targets realtime dialogue / customer-service scenarios — runs on a single consumer GPU. Streaming ASR is sensitive to memory and CPU scheduling — 16 GB+ workstation or cloud inference recommended. Watch community quantization and ONNX exports.