← News
News

Alibaba Launches Qwen3.8-Omni-Flash: Native Omni-Modal, 1M Context, Audio Costs Down 98%

Alibaba's Qwen team has released Qwen3.8-Omni-Flash: a natively omni-modal model built on the Qwen3.8-Flash-Next architecture, accepting text/image/audio/video input with a 1M-token context at $0.15 per million input tokens. Hourly audio input cost drops over 98% versus its predecessor. API-only; no open weights.

Source: DataCamp

Launch

On September 18, 2026, Alibaba's Qwen team released Qwen3.8-Omni-Flash — a natively omni-modal model built on the Qwen3.8-Flash-Next architecture, positioned at the Flash tier (high cost-efficiency, high throughput).

Modalities and Context

  • Input: native support for text, image, audio, and video; output is text only (no speech/video generation)
  • 1M-token context: up to ~991K input tokens, 131K max output, and a 262K thinking-mode reasoning budget
  • Audio covers 113 languages and dialects, with dual- and quad-channel spatial audio understanding; speaker diarization error in meeting scenarios dropped from 88% to 3%

Performance and Cost

  • Qwen reports a 25%+ average improvement across 29 evaluations versus the previous Qwen3.5-Omni-Plus
  • Agentic mode improved on OmniVideoBench from 63.4 to 67.8 while cutting token consumption by ~45.7%
  • API pricing: $0.15/M input tokens ($0.016 cached), $0.47/M output tokens
  • Hourly audio input cost is down over 98% versus the previous generation; joint audio-video input down over 93%

Open-Source Companions and Access

  • Companion open-source releases: Qwen-MM-Plugins (agent plugin suite) and Qwen-Live Harness (realtime interaction runtime for the Realtime variant)
  • Available via Alibaba Cloud Model Studio (DashScope- and OpenAI-compatible), live in Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, Virginia, and more
  • API-only — no open weights. For self-hosters, see open-weight options we track such as Qwen3.8-Flash-Next