← News
News

Qwen3.8-Flash-Next architecture deep dive: 48-layer hybrid (Gated DeltaNet + Sparse Attention), 6B active / 180B total

Qwen3.8-Flash-Next breakdown: 48 layers (36 Gated DeltaNet + 12 Qwen Sparse Attention), 512 routed experts top-10 + 1 shared, native 262K ctx (YaRN to 1M), 180B total / 6B active, llama.cpp codename qwen4exp (Qwen4 architecture preview).

Source: RunAIHome + HuggingFace ISTA-DASLab

What it is

Qwen3.8-Flash-Next is the "Flash tier" flagship MoE released by Alibaba's Qwen team on 2026-08-26: 180B total, only 6B active per token. It was designed for "runnable on consumer hardware" as a primary goal, which led to a rare hybrid architecture.

Key parameters

Dimension Value
Total params ~180B (125B main + 51B n-gram embeddings + 4B MTP head)
Active params ~6B / token
Layers 48
Attention 36 layers Gated DeltaNet + 12 layers Qwen Sparse Attention
Routed experts 512, top-10 activated + 1 shared
Context native 262K, YaRN to 1M
Multimodal vision + text
License Qwen Community License 1.0
llama.cpp codename qwen4exp (Qwen4 architecture preview)

Why "6B active" is the key

The 180B total can sit on SSD/RAM without issue — but only 6B is computed per token, so VRAM pressure drops from "fit all weights" to "fit 6B active + KV cache". This is the prerequisite that makes tiered-memory solutions like Strata viable.

Cost of DeltaNet + Sparse Attention hybrid

The 36 Gated DeltaNet layers use linear attention (O(n) instead of O(n²)), keeping long-context decode fast; the 12 Sparse Attention layers preserve full attention for tokens that need precise alignment (code, math). The benefit is stable decode speed at 1M context; the cost is implementation complexity, and frameworks lag behind pure-Transformer support.

Comparison (180B tier)

Model Total Active Context Architecture
Qwen3.8-Flash-Next 180B 6B 262K (1M YaRN) DeltaNet + Sparse Attention MoE
DeepSeek-V4.1-Flash 256B ~8B 1M Dense-MoE + MLA
GLM-5.3-Flash 200B ~6B 200K Dense-MoE

Local deployment angle

For this model, framework priority: Strata > llama.cpp (latest main supports qwen4exp) > vLLM (SGLang hasn't caught up yet). For quantization, IQ2_XS is the default (45GB weights + 64GB RAM); for quality-sensitive scenarios, IQ3_XXS (76GB, needs 96GB RAM).