[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f3jjmhf6g5i5sm":3},{"items":4,"total":75,"page":107,"pageSize":108},[5,18,30,39,48,57,66,78,87,97],{"slug":6,"category":7,"publishedAt":8,"titleZh":9,"titleEn":10,"summaryZh":11,"summaryEn":12,"models":13,"hardwares":17},"kimi-k3-glm-5-3-openai-enterprise","NEWS","2026-10-02T05:15:00.000Z","中国开源模型首次进入 OpenAI 企业结算体系：Kimi K3、GLM 5.3 接入 Codex 企业通道","Chinese Open Models Enter OpenAI Enterprise Billing for the First Time: Kimi K3 and GLM 5.3 Join the Codex Enterprise Channel","据多家中文媒体报道，Baseten 成为 OpenAI B2B 市场首批开源模型推理服务商：企业可在 Codex \u002F Responses API 中调用 Kimi K3、GLM 5.3，费用计入既有 OpenAI 采购承诺额度——中国开源模型首次进入 OpenAI 企业付费结算体系。","According to multiple Chinese media reports, Baseten is among the first open-model inference providers in OpenAI's B2B marketplace: enterprises can call Kimi K3 and GLM 5.3 through Codex or the Responses API, billed against their existing OpenAI procurement commitments — the first time Chinese open models enter OpenAI enterprise billing.",[14],{"id":15,"name":16},20,"GLM-5.3",[],{"slug":19,"category":7,"publishedAt":20,"titleZh":21,"titleEn":22,"summaryZh":23,"summaryEn":24,"models":25,"hardwares":29},"deepseek-v4-1-flash-derivatives","2026-10-02T05:10:00.000Z","发布三周，DeepSeek V4.1 Flash 衍生生态成形：GGUF 量化、abliterated 与 FP8 社区版","Three Weeks On, the DeepSeek V4.1 Flash Derivative Ecosystem Takes Shape: GGUF Quants, Abliterated and FP8 Community Builds","得益于 MIT 许可，DeepSeek V4.1 Flash 发布数小时内社区衍生版即上线：abliterated 全精度版、GGUF 本地量化版、FP8 uncensored 版（下载量最高）相继出现，另有 MixedQ2 混合量化与 forgequant 量化工具。官方未参与或背书任何衍生版本。","Thanks to the MIT license, community derivatives of DeepSeek V4.1 Flash appeared within hours of release: a full-precision abliterated build, GGUF quants for local inference, and an FP8 uncensored variant (the most downloaded), plus a MixedQ2 mixed-quant build and the forgequant quantization toolkit. DeepSeek has not created or endorsed any derivative.",[26],{"id":27,"name":28},18,"DeepSeek-V4.1-Flash",[],{"slug":31,"category":7,"publishedAt":32,"titleZh":33,"titleEn":34,"summaryZh":35,"summaryEn":36,"models":37,"hardwares":38},"ling-3-1-flash","2026-10-02T05:05:00.000Z","蚂蚁百灵发布 Ling-3.1-flash：560B 总参数 \u002F 25B 激活，连续 17 小时写出编译器","Ant's InclusionAI Ships Ling-3.1-flash: 560B Total \u002F 25B Active, Wrote a Compiler in 17 Straight Hours","蚂蚁百灵（InclusionAI）发布 Ling-3.1-flash：约 560B 总参数、25B 激活，混合线性注意力架构，1M 上下文（体验期 256K）；官方演示连续 17 小时从零写出 Lua 到 x86-64 编译器（182 项测试过 178）。目前未开源，官方称转付费后计划开源。","Ant Group's InclusionAI has released Ling-3.1-flash: ~560B total parameters with ~25B active, a hybrid linear-attention architecture, and a 1M-token context (256K during the free trial). Official demos show it writing a Lua-to-x86-64 compiler from scratch over 17 straight hours (178\u002F182 tests passed). Not yet open-sourced — the weights are planned for release when the paid tier opens.",[],[],{"slug":40,"category":7,"publishedAt":41,"titleZh":42,"titleEn":43,"summaryZh":44,"summaryEn":45,"models":46,"hardwares":47},"gpt-6-1-sol-devday-2026","2026-10-02T05:00:00.000Z","OpenAI DevDay 2026：GPT-6.1 Sol 发布——逼近旗舰实力，价格五分之一","OpenAI DevDay 2026: GPT-6.1 Sol — Near-Flagship Capability at One-Fifth the Price","OpenAI 在 DevDay 2026 发布 GPT-6.1 Sol（gpt-6.1-sol）：输入 $2\u002F输出 $10 每百万 token，约为 GPT-6 Astra 的五分之一，1.05M 上下文，已成 Codex 默认模型；同场发布 dots 智能体、Pro 500 套餐与 Agents API。此前一天，GPT-6.1 Astra 因内部对齐测试未通过被搁置。","At DevDay 2026, OpenAI launched GPT-6.1 Sol (gpt-6.1-sol): $2\u002F$10 per million tokens — about one-fifth of GPT-6 Astra — with a 1.05M context window, now the default model in Codex. Also announced: dots agents, the Pro 500 plan, and the Agents API. A day earlier, GPT-6.1 Astra was shelved after failing internal alignment tests.",[],[],{"slug":49,"category":7,"publishedAt":50,"titleZh":51,"titleEn":52,"summaryZh":53,"summaryEn":54,"models":55,"hardwares":56},"gemini-4-argon-staged-rollout","2026-10-01T05:20:00.000Z","谷歌发布 Gemini 4（Argon）：安全优先的分阶段开放","Google Unveils Gemini 4 (Argon): A Safety-First Staged Rollout","谷歌发布新一代旗舰 Gemini 4 Argon，主打复杂长周期工作流中的持续深度推理；采取分阶段发布策略，初期仅向可信赖的网络安全防御人员开放，评测与定价均未披露。","Google has unveiled Gemini 4 Argon, its next-generation flagship focused on sustained deep reasoning in complex, long-horizon workflows. The staged rollout opens first to trusted cyber-defense personnel; benchmarks and pricing remain undisclosed.",[],[],{"slug":58,"category":7,"publishedAt":59,"titleZh":60,"titleEn":61,"summaryZh":62,"summaryEn":63,"models":64,"hardwares":65},"claude-opus-5-5","2026-10-01T05:15:00.000Z","Anthropic 发布 Claude Opus 5.5：$4\u002F$20 定价，旗舰门槛再降","Anthropic Ships Claude Opus 5.5: $4\u002F$20 Pricing Lowers the Flagship Bar","Anthropic 发布 Claude 5.5 系列首款模型 Opus 5.5（claude-opus-5-5）：输入 $4\u002F输出 $20 每百万 token，较 Opus 5 降价 20%，1M 上下文，官方称多数工作负载达到 Fable 5.1 水平而价格不到一半；Sonnet 5.5 六天后跟进。","Anthropic has released Opus 5.5 (claude-opus-5-5), the first model of the Claude 5.5 family: $4\u002F$20 per million tokens — 20% below Opus 5 — with a 1M-token context window. Anthropic claims Fable 5.1-level results on most work at under half the price. Sonnet 5.5 followed six days later.",[],[],{"slug":67,"category":7,"publishedAt":68,"titleZh":69,"titleEn":70,"summaryZh":71,"summaryEn":72,"models":73,"hardwares":77},"qwen3-8-omni-flash","2026-10-01T05:10:00.000Z","阿里发布 Qwen3.8-Omni-Flash：1M 上下文原生全模态，音频成本降 98%","Alibaba Launches Qwen3.8-Omni-Flash: Native Omni-Modal, 1M Context, Audio Costs Down 98%","阿里千问发布 Qwen3.8-Omni-Flash：基于 Qwen3.8-Flash-Next 架构的原生全模态模型，支持文本\u002F图像\u002F音频\u002F视频输入与 1M 上下文，输入 $0.15\u002F百万 token，每小时音频输入成本较上代降超 98%；仅 API 提供，未开放权重。","Alibaba's Qwen team has released Qwen3.8-Omni-Flash: a natively omni-modal model built on the Qwen3.8-Flash-Next architecture, accepting text\u002Fimage\u002Faudio\u002Fvideo input with a 1M-token context at $0.15 per million input tokens. Hourly audio input cost drops over 98% versus its predecessor. API-only; no open weights.",[74],{"id":75,"name":76},11,"Qwen3.8-Flash-Next",[],{"slug":79,"category":7,"publishedAt":80,"titleZh":81,"titleEn":82,"summaryZh":83,"summaryEn":84,"models":85,"hardwares":86},"xiaomi-mimo-v2-6-open","2026-10-01T05:05:00.000Z","小米开源 MiMo-V2.6：Pro 版登顶 Artificial Analysis 开放权重榜","Xiaomi Open-Sources MiMo-V2.6: Pro Tops the Artificial Analysis Open-Weight Board","小米发布并开源 MiMo-V2.6 系列（完全开放权重）：分 Pro \u002F Flash \u002F Ultraspeed 三档并附桌面客户端，全系原生全模态；据行业日报报道，Pro 版登顶 Artificial Analysis 开放权重模型榜，分数高于 Kimi K3 与 Qwen3.8 Max。","Xiaomi has released and open-sourced the MiMo-V2.6 series with fully open weights: three tiers (Pro \u002F Flash \u002F Ultraspeed) plus a desktop client, all natively omni-modal. Industry digests report the Pro tier tops the Artificial Analysis open-weight leaderboard, ahead of Kimi K3 and Qwen3.8 Max.",[],[],{"slug":88,"category":7,"publishedAt":89,"titleZh":90,"titleEn":91,"summaryZh":92,"summaryEn":93,"models":94,"hardwares":96},"deepseek-v4-1-flash-open","2026-10-01T05:00:00.000Z","DeepSeek V4.1 Flash 开源：每 token KV Cache 仅 890 字节，百万上下文不再昂贵","DeepSeek V4.1 Flash Open-Sourced: 890 Bytes of KV Cache per Token Makes 1M Context Cheap","DeepSeek 发布并开源 V4.1 Flash（MIT 许可）：552B 非对称 MoE，每 token 全局 KV Cache 仅 890 字节、约为上代 1\u002F4，1M 上下文 KV 仅约 890MB；Agent 基准超越 V4 Pro，API 输入 $0.30\u002F百万 token、闲时半价。","DeepSeek has released and open-sourced V4.1 Flash under MIT: a 552B asymmetric MoE whose global KV cache is just 890 bytes per token — a quarter of its predecessor — bringing a 1M-token context down to ~890MB of KV. It beats V4 Pro on agent benchmarks; API input is $0.30 per million tokens with half-price off-peak.",[95],{"id":27,"name":28},[],{"slug":98,"category":99,"publishedAt":100,"titleZh":101,"titleEn":102,"summaryZh":103,"summaryEn":104,"models":105,"hardwares":106},"llamacpp-manual","GUIDE","2026-09-20T00:00:00.000Z","llama.cpp 手动部署","Manual setup with llama.cpp","下载 GGUF、自选量化与参数，理解每一步在做什么——可控性与性能天花板都更高。","Download GGUF weights, pick your own quantization and flags, and understand every step — more control, higher performance ceiling.",[],[],1,10]