News

What is changing in local AI — model releases, hardware news, runtime updates, and deployment guides.

News

News

Chinese Open Models Enter OpenAI Enterprise Billing for the First Time: Kimi K3 and GLM 5.3 Join the Codex Enterprise Channel

According to multiple Chinese media reports, Baseten is among the first open-model inference providers in OpenAI's B2B marketplace: enterprises can call Kimi K3 and GLM 5.3 through Codex or the Responses API, billed against their existing OpenAI procurement commitments — the first time Chinese open models enter OpenAI enterprise billing.

GLM-5.3
News

Three Weeks On, the DeepSeek V4.1 Flash Derivative Ecosystem Takes Shape: GGUF Quants, Abliterated and FP8 Community Builds

Thanks to the MIT license, community derivatives of DeepSeek V4.1 Flash appeared within hours of release: a full-precision abliterated build, GGUF quants for local inference, and an FP8 uncensored variant (the most downloaded), plus a MixedQ2 mixed-quant build and the forgequant quantization toolkit. DeepSeek has not created or endorsed any derivative.

DeepSeek-V4.1-Flash
News

Ant's InclusionAI Ships Ling-3.1-flash: 560B Total / 25B Active, Wrote a Compiler in 17 Straight Hours

Ant Group's InclusionAI has released Ling-3.1-flash: ~560B total parameters with ~25B active, a hybrid linear-attention architecture, and a 1M-token context (256K during the free trial). Official demos show it writing a Lua-to-x86-64 compiler from scratch over 17 straight hours (178/182 tests passed). Not yet open-sourced — the weights are planned for release when the paid tier opens.

News

OpenAI DevDay 2026: GPT-6.1 Sol — Near-Flagship Capability at One-Fifth the Price

At DevDay 2026, OpenAI launched GPT-6.1 Sol (gpt-6.1-sol): $2/$10 per million tokens — about one-fifth of GPT-6 Astra — with a 1.05M context window, now the default model in Codex. Also announced: dots agents, the Pro 500 plan, and the Agents API. A day earlier, GPT-6.1 Astra was shelved after failing internal alignment tests.

News

Google Unveils Gemini 4 (Argon): A Safety-First Staged Rollout

Google has unveiled Gemini 4 Argon, its next-generation flagship focused on sustained deep reasoning in complex, long-horizon workflows. The staged rollout opens first to trusted cyber-defense personnel; benchmarks and pricing remain undisclosed.

News

Anthropic Ships Claude Opus 5.5: $4/$20 Pricing Lowers the Flagship Bar

Anthropic has released Opus 5.5 (claude-opus-5-5), the first model of the Claude 5.5 family: $4/$20 per million tokens — 20% below Opus 5 — with a 1M-token context window. Anthropic claims Fable 5.1-level results on most work at under half the price. Sonnet 5.5 followed six days later.

News

Alibaba Launches Qwen3.8-Omni-Flash: Native Omni-Modal, 1M Context, Audio Costs Down 98%

Alibaba's Qwen team has released Qwen3.8-Omni-Flash: a natively omni-modal model built on the Qwen3.8-Flash-Next architecture, accepting text/image/audio/video input with a 1M-token context at $0.15 per million input tokens. Hourly audio input cost drops over 98% versus its predecessor. API-only; no open weights.

Qwen3.8-Flash-Next
News

Xiaomi Open-Sources MiMo-V2.6: Pro Tops the Artificial Analysis Open-Weight Board

Xiaomi has released and open-sourced the MiMo-V2.6 series with fully open weights: three tiers (Pro / Flash / Ultraspeed) plus a desktop client, all natively omni-modal. Industry digests report the Pro tier tops the Artificial Analysis open-weight leaderboard, ahead of Kimi K3 and Qwen3.8 Max.

News

DeepSeek V4.1 Flash Open-Sourced: 890 Bytes of KV Cache per Token Makes 1M Context Cheap

DeepSeek has released and open-sourced V4.1 Flash under MIT: a 552B asymmetric MoE whose global KV cache is just 890 bytes per token — a quarter of its predecessor — bringing a 1M-token context down to ~890MB of KV. It beats V4 Pro on agent benchmarks; API input is $0.30 per million tokens with half-price off-peak.

DeepSeek-V4.1-Flash
Guides

Manual setup with llama.cpp

Download GGUF weights, pick your own quantization and flags, and understand every step — more control, higher performance ceiling.

11 articles