[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f1lph1y6wjppro":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":15,"bodyZh":16,"bodyEn":17,"sourceName":18,"sourceUrl":19,"updatedAt":20},"qwen3-8-flash-next-architecture","NEWS","2026-10-05T04:24:00.000Z","Qwen3.8-Flash-Next 架构深度：48 层混合（Gated DeltaNet + Sparse Attention）+ 6B 激活 \u002F 180B 总","Qwen3.8-Flash-Next architecture deep dive: 48-layer hybrid (Gated DeltaNet + Sparse Attention), 6B active \u002F 180B total","拆解 Qwen3.8-Flash-Next：48 层（36 Gated DeltaNet + 12 Qwen Sparse Attention）、512 路由专家 top-10 + 1 共享、原生 262K 上下文（YaRN 扩到 1M）、180B 总参数仅 6B 激活、llama.cpp 内部代号 qwen4exp（Qwen4 架构预览）。","Qwen3.8-Flash-Next breakdown: 48 layers (36 Gated DeltaNet + 12 Qwen Sparse Attention), 512 routed experts top-10 + 1 shared, native 262K ctx (YaRN to 1M), 180B total \u002F 6B active, llama.cpp codename qwen4exp (Qwen4 architecture preview).",[12],{"id":13,"name":14},11,"Qwen3.8-Flash-Next",[],"### 是什么\n\nQwen3.8-Flash-Next 是阿里 Qwen 团队 2026 年 8 月 26 日发布的\"Flash 档\"旗舰 MoE——**180B 总参数、仅 6B 激活**（每个 token）。它把\"在消费级硬件上跑得动\"作为首要目标，于是采用了一套非常少见的混合架构。\n\n### 关键参数\n\n| 维度 | 数值 |\n|---|---|\n| 总参数 | 约 180B（125B 主模型 + 51B n-gram 嵌入 + 4B MTP head） |\n| 激活参数 | 约 6B \u002F token |\n| 层数 | 48 层 |\n| 注意力机制 | 36 层 **Gated DeltaNet** + 12 层 **Qwen Sparse Attention** |\n| 路由专家 | 512 个，top-10 激活 + 1 个共享 |\n| 上下文 | 原生 262K，YaRN 可扩到 1M |\n| 多模态 | 视觉 + 文本 |\n| 许可证 | Qwen Community License 1.0 |\n| llama.cpp 代号 | **qwen4exp**（Qwen4 架构预览） |\n\n### 为什么\"6B 激活\"是关键\n\n180B 总参数放 SSD\u002FRAM 上没问题——但**每个 token 真正计算只涉及 6B**，于是 GPU 显存压力瞬间从\"装下全部权重\"降到\"装下 6B 激活 + KV cache\"。这是 Strata 等分层内存方案能跑通的核心前置。\n\n### DeltaNet + Sparse Attention 混合的代价\n\n36 层 Gated DeltaNet 是线性注意力（O(n) 而非 O(n²)），能极快处理长上下文；12 层 Sparse Attention 保留全注意力，处理需要精确对齐的 token（比如代码、数学）。混合的好处是 1M 上下文也能维持 decode 速度不崩；坏处是工程实现复杂，框架支持滞后于纯 Transformer。\n\n### 与同类对比（180B 档）\n\n| 模型 | 总参数 | 激活 | 上下文 | 架构 |\n|---|---|---|---|---|\n| Qwen3.8-Flash-Next | 180B | 6B | 262K (1M YaRN) | DeltaNet + Sparse Attention MoE |\n| DeepSeek-V4.1-Flash | 256B | ~8B | 1M | Dense-MoE + MLA |\n| GLM-5.3-Flash | 200B | ~6B | 200K | Dense-MoE |\n\n### 本地部署角度\n\n跑这个模型，**框架选择优先级**：Strata > llama.cpp（最新 main 分支持 qwen4exp）> vLLM（SGLang 暂未跟进）。量化首选 IQ2_XS（45GB 权重 + 64GB RAM 装下），质量敏感场景用 IQ3_XXS（76GB，需 96GB RAM）。","### What it is\n\nQwen3.8-Flash-Next is the \"Flash tier\" flagship MoE released by Alibaba's Qwen team on 2026-08-26: **180B total, only 6B active per token**. It was designed for \"runnable on consumer hardware\" as a primary goal, which led to a rare hybrid architecture.\n\n### Key parameters\n\n| Dimension | Value |\n|---|---|\n| Total params | ~180B (125B main + 51B n-gram embeddings + 4B MTP head) |\n| Active params | ~6B \u002F token |\n| Layers | 48 |\n| Attention | 36 layers **Gated DeltaNet** + 12 layers **Qwen Sparse Attention** |\n| Routed experts | 512, top-10 activated + 1 shared |\n| Context | native 262K, YaRN to 1M |\n| Multimodal | vision + text |\n| License | Qwen Community License 1.0 |\n| llama.cpp codename | **qwen4exp** (Qwen4 architecture preview) |\n\n### Why \"6B active\" is the key\n\nThe 180B total can sit on SSD\u002FRAM without issue — but **only 6B is computed per token**, so VRAM pressure drops from \"fit all weights\" to \"fit 6B active + KV cache\". This is the prerequisite that makes tiered-memory solutions like Strata viable.\n\n### Cost of DeltaNet + Sparse Attention hybrid\n\nThe 36 Gated DeltaNet layers use linear attention (O(n) instead of O(n²)), keeping long-context decode fast; the 12 Sparse Attention layers preserve full attention for tokens that need precise alignment (code, math). The benefit is stable decode speed at 1M context; the cost is implementation complexity, and frameworks lag behind pure-Transformer support.\n\n### Comparison (180B tier)\n\n| Model | Total | Active | Context | Architecture |\n|---|---|---|---|---|\n| Qwen3.8-Flash-Next | 180B | 6B | 262K (1M YaRN) | DeltaNet + Sparse Attention MoE |\n| DeepSeek-V4.1-Flash | 256B | ~8B | 1M | Dense-MoE + MLA |\n| GLM-5.3-Flash | 200B | ~6B | 200K | Dense-MoE |\n\n### Local deployment angle\n\nFor this model, **framework priority**: Strata > llama.cpp (latest main supports qwen4exp) > vLLM (SGLang hasn't caught up yet). For quantization, IQ2_XS is the default (45GB weights + 64GB RAM); for quality-sensitive scenarios, IQ3_XXS (76GB, needs 96GB RAM).","RunAIHome + HuggingFace ISTA-DASLab","https:\u002F\u002Frunaihome.com\u002Fblog\u002Fqwen38-flash-next-local-ai-hardware-guide-2026","2026-10-05T02:23:19.778Z"]