DeepSeek V4.1 Flash Open-Sourced: 890 Bytes of KV Cache per Token Makes 1M Context Cheap
DeepSeek has released and open-sourced V4.1 Flash under MIT: a 552B asymmetric MoE whose global KV cache is just 890 bytes per token — a quarter of its predecessor — bringing a 1M-token context down to ~890MB of KV. It beats V4 Pro on agent benchmarks; API input is $0.30 per million tokens with half-price off-peak.
Source: Hugging Face(deepseek-ai)
Launch and Open Source
On September 10, 2026, DeepSeek released and open-sourced DeepSeek V4.1 Flash under the MIT license, with weights on Hugging Face and ModelScope (deepseek-ai/DeepSeek-V4.1-Flash, 48 shards), accompanied by a 51-page technical report Pushing the Limits of KV Cache Compression, a prompt-encoding reference, and a minimal PyTorch inference implementation. The same day also brought the DeepJIT kernel JIT library and the Harness v0.1.5 agent framework.
The Core Story: KV Cache Pushed to the Limit
- Just 890 bytes of global KV cache per token — a quarter of the previous V4 Flash (3,514 B), and 437× tighter than the original generation
- A 1M-token context needs only ~890MB of global KV (versus 48GB for V3.2 on the same footing)
- Key techniques: three CSA2 layer modes, FP4 caching, bounded replay with sliding-window attention, cross-layer attention reuse
- HBM requirements drop to 1/4 of the predecessor; SSD requirements to 1/8
For long-context agent sessions this is the most direct cost cut — KV cache is exactly where the memory bill of long sessions lives.
Architecture
- Causal Encoder-Decoder (CED) asymmetric MoE: 20 encoder layers + 20 decoder layers
- 552B total parameters (plus 196B Engram parameters and a ViT vision encoder; native image input)
- Activated parameters: just 8B for prefill, 16B for decode
- 384 routed experts + 1 shared expert, 6 active per token, FP4 expert weights
- 1M-token context (max position embedding 1,048,576), 384K max output
Performance and Pricing
- Agent benchmarks: DeepSWE v1.1 at 74.2%, Terminal-Bench 2.1 at 90.6%, ahead of V4 Pro
- Topped the Artificial Analysis open-weight board at launch (Intelligence Index 40)
- Unified API name
deepseek-flash: $0.30/M input, $1.20/M output, half price off-peak, cache-hit input as low as $0.006/M - V4 Pro has been routed to V4.1 Flash at Flash pricing since September 14
What It Means for Self-Hosting
The official self-hosting bar — roughly 2,000 GPUs plus a storage cluster — remains out of reach; but the community has already run it on 4× DGX Spark with vLLM. A 4× smaller KV cache means much longer contexts within the same VRAM — exactly what this site tracks. We list DeepSeek-V4.1-Flash; watch for its quantized releases.