Ant's InclusionAI Ships Ling-3.1-flash: 560B Total / 25B Active, Wrote a Compiler in 17 Straight Hours
Ant Group's InclusionAI has released Ling-3.1-flash: ~560B total parameters with ~25B active, a hybrid linear-attention architecture, and a 1M-token context (256K during the free trial). Official demos show it writing a Lua-to-x86-64 compiler from scratch over 17 straight hours (178/182 tests passed). Not yet open-sourced — the weights are planned for release when the paid tier opens.
Source: The BlockBeats
Launch
On September 30, 2026, Ant Group's InclusionAI released Ling-3.1-flash. Compared with the previous Ling-3.0-flash (124B total / 5.1B active), the total size grows about 4.5×.
Architecture and Specs
- ~560B total parameters, ~25B active per token (sparsity ratio ~1:22)
- Continues the hybrid linear-attention design with a higher share of linear layers: 7 KDA layers + 1 Gated MLA
- 512 routed experts, 8 active per token + 1 shared expert
- Up to 1M-token context (256K during the trial; 1M unlocks with the paid tier)
Official Demos (Self-Reported)
- ~17 straight hours writing a Lua-to-x86-64 ELF compiler from scratch: 178 of 182 tests passed (97.8%)
- ~20 hours porting a C image library to Rust: 8.015× speedup, all 30 correctness checks passed
- Benchmarks: GDPVal-AA v2.1 at 1,673 Elo, FrontierSWE 75.16, HealthBench Professional 65.35
Open-Source Status: Not Yet Open
- Official line: when the two-week free trial converts to paid, the 1M context opens and the model is planned to be open-sourced alongside
- As of October 1, there is no weights repository, license, or model card on Hugging Face; all scores are self-reported with no third-party evaluation
- The previous Ling-3.0-flash took about two weeks from launch to open weights; expect a similar cadence
Where to Try
A two-week free trial is live via Novita AI, Vercel AI Gateway, and Ant's own Ling Studio (ling.tbox.cn), positioned for general agents, search, office work, and software development, plus medical, finance, and materials-science scenarios.
If the weights land as planned, the 560B/25B sparsity means decode-time memory pressure concentrates on the 25B active side — an interesting new option for local multi-GPU or large-VRAM single-card deployments. We'll keep watching.