[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f2107f5z57o30k":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":15,"bodyZh":16,"bodyEn":17,"sourceName":18,"sourceUrl":19,"updatedAt":20},"qwen3-8-omni-flash","NEWS","2026-10-01T05:10:00.000Z","阿里发布 Qwen3.8-Omni-Flash：1M 上下文原生全模态，音频成本降 98%","Alibaba Launches Qwen3.8-Omni-Flash: Native Omni-Modal, 1M Context, Audio Costs Down 98%","阿里千问发布 Qwen3.8-Omni-Flash：基于 Qwen3.8-Flash-Next 架构的原生全模态模型，支持文本\u002F图像\u002F音频\u002F视频输入与 1M 上下文，输入 $0.15\u002F百万 token，每小时音频输入成本较上代降超 98%；仅 API 提供，未开放权重。","Alibaba's Qwen team has released Qwen3.8-Omni-Flash: a natively omni-modal model built on the Qwen3.8-Flash-Next architecture, accepting text\u002Fimage\u002Faudio\u002Fvideo input with a 1M-token context at $0.15 per million input tokens. Hourly audio input cost drops over 98% versus its predecessor. API-only; no open weights.",[12],{"id":13,"name":14},11,"Qwen3.8-Flash-Next",[],"## 发布\n\n2026 年 9 月 18 日，阿里千问团队发布 **Qwen3.8-Omni-Flash**——基于 **Qwen3.8-Flash-Next** 架构的原生全模态模型，定位 Flash 层级（高性价比、高吞吐）。\n\n## 模态与上下文\n\n- 输入：文本、图像、音频、视频四类原生支持；输出为文本（不做语音\u002F视频生成）\n- **1M token 上下文**：最大输入约 99.1 万 token，最大输出 13.1 万，思考模式推理预算 262K\n- 音频支持 **113 种语言与方言**，双声道与 4 声道空间音频理解；会议场景说话人分离错误率从 88% 降至 3%\n\n## 性能与成本\n\n- 官方称 29 项评测平均分较上代 Qwen3.5-Omni-Plus 提升超 25%\n- Agentic 模式在 OmniVideoBench 上从 63.4 提升至 67.8，同时 token 消耗降低约 45.7%\n- API 定价：输入 **$0.15**\u002F百万 token（缓存输入 $0.016），输出 **$0.47**\u002F百万 token\n- 每小时音频输入成本较上代下降超 **98%**，音视频联合输入下降超 93%\n\n## 配套开源与接入\n\n- 配套开源 **Qwen-MM-Plugins**（智能体插件套件）与 **Qwen-Live Harness**（实时交互运行时，配合 Realtime 版本）\n- 通过阿里云百炼提供 API（兼容 DashScope 与 OpenAI 协议），已上线北京、新加坡、香港、东京、法兰克福、弗吉尼亚等区域\n- **仅 API 提供，未发布开放权重**——对本地部署玩家而言，可继续关注我们收录的 [Qwen3.8-Flash-Next](\u002Fmodel\u002FQwen3.8-Flash-Next) 等开放权重型号","## Launch\n\nOn September 18, 2026, Alibaba's Qwen team released **Qwen3.8-Omni-Flash** — a natively omni-modal model built on the **Qwen3.8-Flash-Next** architecture, positioned at the Flash tier (high cost-efficiency, high throughput).\n\n## Modalities and Context\n\n- Input: native support for text, image, audio, and video; output is text only (no speech\u002Fvideo generation)\n- **1M-token context**: up to ~991K input tokens, 131K max output, and a 262K thinking-mode reasoning budget\n- Audio covers **113 languages and dialects**, with dual- and quad-channel spatial audio understanding; speaker diarization error in meeting scenarios dropped from 88% to 3%\n\n## Performance and Cost\n\n- Qwen reports a 25%+ average improvement across 29 evaluations versus the previous Qwen3.5-Omni-Plus\n- Agentic mode improved on OmniVideoBench from 63.4 to 67.8 while cutting token consumption by ~45.7%\n- API pricing: **$0.15**\u002FM input tokens ($0.016 cached), **$0.47**\u002FM output tokens\n- Hourly audio input cost is down over **98%** versus the previous generation; joint audio-video input down over 93%\n\n## Open-Source Companions and Access\n\n- Companion open-source releases: **Qwen-MM-Plugins** (agent plugin suite) and **Qwen-Live Harness** (realtime interaction runtime for the Realtime variant)\n- Available via Alibaba Cloud Model Studio (DashScope- and OpenAI-compatible), live in Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, Virginia, and more\n- **API-only — no open weights.** For self-hosters, see open-weight options we track such as [Qwen3.8-Flash-Next](\u002Fen\u002Fmodel\u002FQwen3.8-Flash-Next)","DataCamp","https:\u002F\u002Fwww.datacamp.com\u002Fblog\u002Fqwen3-8-omni-flash","2026-10-01T05:14:10.750Z"]