Qwen3.8-Flash-Next architecture deep dive: 48-layer hybrid (Gated DeltaNet + Sparse Attention), 6B active / 180B total
Qwen3.8-Flash-Next breakdown: 48 layers (36 Gated DeltaNet + 12 Qwen Sparse Attention), 512 routed experts top-10 + 1 shared, native 262K ctx (YaRN to 1M), 180B total / 6B active, llama.cpp codename qwen4exp (Qwen4 architecture preview).
What it is
Qwen3.8-Flash-Next is the "Flash tier" flagship MoE released by Alibaba's Qwen team on 2026-08-26: 180B total, only 6B active per token. It was designed for "runnable on consumer hardware" as a primary goal, which led to a rare hybrid architecture.
Key parameters
| Dimension | Value |
|---|---|
| Total params | ~180B (125B main + 51B n-gram embeddings + 4B MTP head) |
| Active params | ~6B / token |
| Layers | 48 |
| Attention | 36 layers Gated DeltaNet + 12 layers Qwen Sparse Attention |
| Routed experts | 512, top-10 activated + 1 shared |
| Context | native 262K, YaRN to 1M |
| Multimodal | vision + text |
| License | Qwen Community License 1.0 |
| llama.cpp codename | qwen4exp (Qwen4 architecture preview) |
Why "6B active" is the key
The 180B total can sit on SSD/RAM without issue — but only 6B is computed per token, so VRAM pressure drops from "fit all weights" to "fit 6B active + KV cache". This is the prerequisite that makes tiered-memory solutions like Strata viable.
Cost of DeltaNet + Sparse Attention hybrid
The 36 Gated DeltaNet layers use linear attention (O(n) instead of O(n²)), keeping long-context decode fast; the 12 Sparse Attention layers preserve full attention for tokens that need precise alignment (code, math). The benefit is stable decode speed at 1M context; the cost is implementation complexity, and frameworks lag behind pure-Transformer support.
Comparison (180B tier)
| Model | Total | Active | Context | Architecture |
|---|---|---|---|---|
| Qwen3.8-Flash-Next | 180B | 6B | 262K (1M YaRN) | DeltaNet + Sparse Attention MoE |
| DeepSeek-V4.1-Flash | 256B | ~8B | 1M | Dense-MoE + MLA |
| GLM-5.3-Flash | 200B | ~6B | 200K | Dense-MoE |
Local deployment angle
For this model, framework priority: Strata > llama.cpp (latest main supports qwen4exp) > vLLM (SGLang hasn't caught up yet). For quantization, IQ2_XS is the default (45GB weights + 64GB RAM); for quality-sensitive scenarios, IQ3_XXS (76GB, needs 96GB RAM).