[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$ftdotz6a6jjtb":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":12,"bodyZh":13,"bodyEn":14,"sourceName":15,"sourceUrl":16,"updatedAt":17},"microsoft-mai-voice-2-1-streaming","NEWS","2026-10-04T05:12:00.000Z","微软开源 MAI-Voice-2.1 与 MAI-Transcribe-2-Streaming：23\u002F60 语言原生口音，150ms\u002F100ms 延迟","Microsoft open-sources MAI-Voice-2.1 and MAI-Transcribe-2-Streaming: 23\u002F60 native accents, 150ms\u002F100ms latency","微软 MAI 团队 10 月 3 日发布两款开源语音栈模型：MAI-Voice-2.1（23 种语言原生口音，Flash 版 150ms 首包延迟）与 MAI-Transcribe-2-Streaming（60 种语言实时转写，首批延迟 ~100ms）。","Microsoft's MAI team released two open-source speech-stack models on October 3: MAI-Voice-2.1 (23 languages with native accents, Flash variant 150 ms first-packet latency) and MAI-Transcribe-2-Streaming (real-time transcription in 60 languages, ~100 ms first-batch latency).",[],[],"### 是什么\n\n10 月 3 日，微软 MAI（Microsoft AI）团队发布两款开源语音栈模型：\n\n- **MAI-Voice-2.1**：TTS 模型，**23 种语言原生口音**（非通用英文转写方言），Flash 版首包延迟约 **150 毫秒**\n- **MAI-Transcribe-2-Streaming**：实时转写模型，**60 种语言**，首批延迟约 **100 毫秒**，流式输出\n\n### 与既有方案的差异\n\n与 Whisper-large-v3 \u002F CosyVoice 等主流开源语音栈相比：\n\n| 维度 | MAI-Voice-2.1 | CosyVoice-2 | Whisper-large-v3 |\n|---|---|---|---|\n| 实时流式 | ✓（Flash 150ms） | ✓（部分） | ✗（离线为主） |\n| 原生口音 | 23 种 | 中文方言为主 | N\u002FA（仅识别） |\n| 首包延迟 | 150ms | ~200ms | N\u002FA |\n| 许可证 | 开源 | 开源（社区） | MIT |\n\n| 维度 | MAI-Transcribe-2-Streaming | Whisper-large-v3 | Paraformer |\n|---|---|---|---|\n| 流式 | ✓ | ✗ | ✓ |\n| 语种数 | 60 | 99 | 中英 |\n| 首批延迟 | 100ms | N\u002FA | ~300ms |\n\n### 本地部署角度\n\nFlash 版 TTS 面向实时对话\u002F客服场景，单张消费级 GPU 即可；流式转写对内存与 CPU 调度敏感，建议 16GB+ 工作站或云端推理服务。社区量化与 ONNX 导出版待跟踪。","### What it is\n\nOn October 3, Microsoft's MAI team released two open-source speech-stack models:\n\n- **MAI-Voice-2.1**: TTS with **23 languages of native accents** (not generic-English-to-local), Flash variant at **~150 ms** first-packet latency\n- **MAI-Transcribe-2-Streaming**: real-time transcription in **60 languages**, **~100 ms** first-batch latency, streaming output\n\n### Comparison\n\nAgainst Whisper-large-v3 \u002F CosyVoice and other mainstream open speech stacks:\n\n| Dimension | MAI-Voice-2.1 | CosyVoice-2 | Whisper-large-v3 |\n|---|---|---|---|\n| Realtime streaming | ✓ (Flash 150ms) | ✓ (partial) | ✗ (offline mostly) |\n| Native accents | 23 | Chinese dialects mostly | N\u002FA (recognition only) |\n| First-packet latency | 150ms | ~200ms | N\u002FA |\n| License | open-source | open-source (community) | MIT |\n\n| Dimension | MAI-Transcribe-2-Streaming | Whisper-large-v3 | Paraformer |\n|---|---|---|---|\n| Streaming | ✓ | ✗ | ✓ |\n| Languages | 60 | 99 | ZH\u002FEN |\n| First-batch latency | 100ms | N\u002FA | ~300ms |\n\n### Local deployment angle\n\nFlash-variant TTS targets realtime dialogue \u002F customer-service scenarios — runs on a single consumer GPU. Streaming ASR is sensitive to memory and CPU scheduling — 16 GB+ workstation or cloud inference recommended. Watch community quantization and ONNX exports.","AI 日报（ai6s.net）","https:\u002F\u002Fai6s.net\u002F6ac0377c05257b0857147348.html","2026-10-04T01:02:09.882Z"]