[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f27p543dg8shmn":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":15,"bodyZh":16,"bodyEn":17,"sourceName":18,"sourceUrl":19,"updatedAt":20},"deepseek-v4-1-flash-open","NEWS","2026-10-01T05:00:00.000Z","DeepSeek V4.1 Flash 开源：每 token KV Cache 仅 890 字节，百万上下文不再昂贵","DeepSeek V4.1 Flash Open-Sourced: 890 Bytes of KV Cache per Token Makes 1M Context Cheap","DeepSeek 发布并开源 V4.1 Flash（MIT 许可）：552B 非对称 MoE，每 token 全局 KV Cache 仅 890 字节、约为上代 1\u002F4，1M 上下文 KV 仅约 890MB；Agent 基准超越 V4 Pro，API 输入 $0.30\u002F百万 token、闲时半价。","DeepSeek has released and open-sourced V4.1 Flash under MIT: a 552B asymmetric MoE whose global KV cache is just 890 bytes per token — a quarter of its predecessor — bringing a 1M-token context down to ~890MB of KV. It beats V4 Pro on agent benchmarks; API input is $0.30 per million tokens with half-price off-peak.",[12],{"id":13,"name":14},18,"DeepSeek-V4.1-Flash",[],"## 发布与开源\n\n2026 年 9 月 10 日，DeepSeek 发布并开源 **DeepSeek V4.1 Flash**，权重以 **MIT 许可**上架 Hugging Face 与 ModelScope（仓库 `deepseek-ai\u002FDeepSeek-V4.1-Flash`，48 个分片），随附 51 页技术报告《Pushing the Limits of KV Cache Compression》、prompt 编码参考与最小化 PyTorch 推理实现。同日还开源了 DeepJIT 内核 JIT 编译库与 Harness v0.1.5 Agent 框架。\n\n## 核心亮点：把 KV Cache 压到极限\n\n- **每 token 全局 KV Cache 仅 890 字节**——约为上代 V4 Flash（3,514 B）的 **1\u002F4**，相比初代压缩 437 倍\n- **1M 上下文的全局 KV 仅约 890MB**（V3.2 同口径为 48GB）\n- 关键技术：CSA2 三种层模式、FP4 缓存、滑动窗口注意力有界重放、跨层注意力复用\n- HBM 显存需求降至上代 1\u002F4，SSD 需求降至 1\u002F8\n\n对长上下文 Agent 会话而言，这是最直接的降本——KV Cache 正是长会话显存账单的大头。\n\n## 架构规格\n\n- Causal Encoder-Decoder（CED）非对称 MoE：20 层编码器 + 20 层解码器\n- 总参数 **552B**（另有 196B Engram 参数 + ViT 视觉编码器，原生支持图像输入）\n- 激活参数：Prefill 仅 8B，Decode 16B\n- 384 个路由专家 + 1 共享专家，每 token 激活 6 个，FP4 专家权重\n- 上下文 100 万 token（最大位置编码 1,048,576），最大输出 384K\n\n## 性能与定价\n\n- Agent 基准 **DeepSWE v1.1 74.2%**、**Terminal-Bench 2.1 90.6%**，超越 V4 Pro\n- 发布时登顶 Artificial Analysis 开放权重榜（智能指数 40）\n- API 统一为 `deepseek-flash`：输入 **$0.30**\u002F百万 token、输出 **$1.20**\u002F百万 token，**闲时半价**，缓存命中输入低至 $0.006\u002FM\n- V4 Pro 自 9 月 14 日起被路由至 V4.1 Flash 并按 Flash 单价计费\n\n## 对本地部署意味着什么\n\n官方自部署门槛约 2000 张 GPU + 存储集群，仍然高不可攀；但社区已有人在 **4 台 DGX Spark + vLLM** 上跑通轻量方案。KV Cache 压缩到 1\u002F4 意味着同等显存下上下文可以拉得更长——这正是本站关注的方向，我们已收录 [DeepSeek-V4.1-Flash](\u002Fmodel\u002FDeepSeek-V4.1-Flash)，可持续关注其量化发布的落地情况。","## Launch and Open Source\n\nOn September 10, 2026, DeepSeek released and open-sourced **DeepSeek V4.1 Flash** under the **MIT license**, with weights on Hugging Face and ModelScope (`deepseek-ai\u002FDeepSeek-V4.1-Flash`, 48 shards), accompanied by a 51-page technical report *Pushing the Limits of KV Cache Compression*, a prompt-encoding reference, and a minimal PyTorch inference implementation. The same day also brought the DeepJIT kernel JIT library and the Harness v0.1.5 agent framework.\n\n## The Core Story: KV Cache Pushed to the Limit\n\n- **Just 890 bytes of global KV cache per token** — a quarter of the previous V4 Flash (3,514 B), and 437× tighter than the original generation\n- **A 1M-token context needs only ~890MB of global KV** (versus 48GB for V3.2 on the same footing)\n- Key techniques: three CSA2 layer modes, FP4 caching, bounded replay with sliding-window attention, cross-layer attention reuse\n- HBM requirements drop to 1\u002F4 of the predecessor; SSD requirements to 1\u002F8\n\nFor long-context agent sessions this is the most direct cost cut — KV cache is exactly where the memory bill of long sessions lives.\n\n## Architecture\n\n- Causal Encoder-Decoder (CED) asymmetric MoE: 20 encoder layers + 20 decoder layers\n- **552B** total parameters (plus 196B Engram parameters and a ViT vision encoder; native image input)\n- Activated parameters: just 8B for prefill, 16B for decode\n- 384 routed experts + 1 shared expert, 6 active per token, FP4 expert weights\n- 1M-token context (max position embedding 1,048,576), 384K max output\n\n## Performance and Pricing\n\n- Agent benchmarks: **DeepSWE v1.1 at 74.2%**, **Terminal-Bench 2.1 at 90.6%**, ahead of V4 Pro\n- Topped the Artificial Analysis open-weight board at launch (Intelligence Index 40)\n- Unified API name `deepseek-flash`: **$0.30**\u002FM input, **$1.20**\u002FM output, **half price off-peak**, cache-hit input as low as $0.006\u002FM\n- V4 Pro has been routed to V4.1 Flash at Flash pricing since September 14\n\n## What It Means for Self-Hosting\n\nThe official self-hosting bar — roughly 2,000 GPUs plus a storage cluster — remains out of reach; but the community has already run it on **4× DGX Spark with vLLM**. A 4× smaller KV cache means much longer contexts within the same VRAM — exactly what this site tracks. We list [DeepSeek-V4.1-Flash](\u002Fen\u002Fmodel\u002FDeepSeek-V4.1-Flash); watch for its quantized releases.","Hugging Face（deepseek-ai）","https:\u002F\u002Fhuggingface.co\u002Fdeepseek-ai\u002FDeepSeek-V4.1-Flash","2026-10-01T05:14:10.779Z"]