← News
News

DeepSeek V4.1 Flash Open-Sourced: 890 Bytes of KV Cache per Token Makes 1M Context Cheap

DeepSeek has released and open-sourced V4.1 Flash under MIT: a 552B asymmetric MoE whose global KV cache is just 890 bytes per token — a quarter of its predecessor — bringing a 1M-token context down to ~890MB of KV. It beats V4 Pro on agent benchmarks; API input is $0.30 per million tokens with half-price off-peak.

Source: Hugging Face(deepseek-ai)

Launch and Open Source

On September 10, 2026, DeepSeek released and open-sourced DeepSeek V4.1 Flash under the MIT license, with weights on Hugging Face and ModelScope (deepseek-ai/DeepSeek-V4.1-Flash, 48 shards), accompanied by a 51-page technical report Pushing the Limits of KV Cache Compression, a prompt-encoding reference, and a minimal PyTorch inference implementation. The same day also brought the DeepJIT kernel JIT library and the Harness v0.1.5 agent framework.

The Core Story: KV Cache Pushed to the Limit

  • Just 890 bytes of global KV cache per token — a quarter of the previous V4 Flash (3,514 B), and 437× tighter than the original generation
  • A 1M-token context needs only ~890MB of global KV (versus 48GB for V3.2 on the same footing)
  • Key techniques: three CSA2 layer modes, FP4 caching, bounded replay with sliding-window attention, cross-layer attention reuse
  • HBM requirements drop to 1/4 of the predecessor; SSD requirements to 1/8

For long-context agent sessions this is the most direct cost cut — KV cache is exactly where the memory bill of long sessions lives.

Architecture

  • Causal Encoder-Decoder (CED) asymmetric MoE: 20 encoder layers + 20 decoder layers
  • 552B total parameters (plus 196B Engram parameters and a ViT vision encoder; native image input)
  • Activated parameters: just 8B for prefill, 16B for decode
  • 384 routed experts + 1 shared expert, 6 active per token, FP4 expert weights
  • 1M-token context (max position embedding 1,048,576), 384K max output

Performance and Pricing

  • Agent benchmarks: DeepSWE v1.1 at 74.2%, Terminal-Bench 2.1 at 90.6%, ahead of V4 Pro
  • Topped the Artificial Analysis open-weight board at launch (Intelligence Index 40)
  • Unified API name deepseek-flash: $0.30/M input, $1.20/M output, half price off-peak, cache-hit input as low as $0.006/M
  • V4 Pro has been routed to V4.1 Flash at Flash pricing since September 14

What It Means for Self-Hosting

The official self-hosting bar — roughly 2,000 GPUs plus a storage cluster — remains out of reach; but the community has already run it on 4× DGX Spark with vLLM. A 4× smaller KV cache means much longer contexts within the same VRAM — exactly what this site tracks. We list DeepSeek-V4.1-Flash; watch for its quantized releases.